Pith. sign in

REVIEW 3 major objections 3 minor 1 cited by

Deploying Foundation Model Powered Agent Services: A Survey

T0 review · 3 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This survey argues that delivering real-time foundation-model agent services at scale requires a unified deployment stack linking hardware execution, resource management, model compression, agent components, and applications.

desk verdict A useful survey of LLM serving and agent frameworks whose framing as a survey of 'agent-service deployment' overstates how much the cited work is actually about agents; the paper's own lessons section concedes the integration is open. read the letter →

arxiv 2412.13437 v1 pith:62L5C7UK submitted 2024-12-18 cs.DC cs.AI

classification cs.DCcs.AI
keywords foundationmodelsAIagentsedge-cloudcomputingmodelservingsystemscompressiontokenreductionresourceallocationparallelism
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to establish that deploying foundation-model (FM) powered agent services in heterogeneous edge-cloud environments is one problem, not a pile of separate ones, and that a unified five-layer framework is the right way to organize both research and practice. The layers run from low-level execution optimization (computation, memory, communication), through resource allocation and parallelism, to model compression and token reduction, then to agent components and applications. The paper reviews existing work inside each layer and argues that jointly optimizing computational and communication resources across these layers is what makes real-time, high-QoS agent services achievable. A sympathetic reader would take the contribution to be a coherent map of the field that identifies where techniques fit and which gaps matter most.

What carries the argument

The central object is the five-layer framework in the paper's Figure 1: an Execution layer, a Resource layer, a Model layer, an Agent layer, and an Application layer. It functions as both a taxonomy and a compositional claim: low-level inference optimizations, resource-allocation and parallelism strategies, model compression and token reduction, agent capabilities, and batching are treated as mutually dependent design choices within one serving stack. The framework carries the argument by showing where each surveyed technique sits and by exposing the missing elasticity at the agent layer as the binding constraint on real-time agent services.

What would settle it

A concrete check would be to search the literature for a prior survey that already covers real-time FM-powered agent deployment across heterogeneous edge-cloud devices under a unified framework; if one exists, the paper's first-comprehensive claim fails. A second check would be an end-to-end experiment combining representative techniques from all five layers (say token reduction, pipeline parallelism, and agent tool calling) to see whether their benefits add up or interfere; the paper reports no such experiment.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that it provides the first comprehensive survey of real-time FM-powered agent service deployment across heterogeneous devices, and that this deployment is best understood through a unified framework of five stacked layers: execution, resource, model, agent, and application. Each lower layer supplies capabilities to the one above it: execution-layer optimizations make inference feasible on diverse hardware; resource-layer parallelism and scaling make the system elastic; model-layer compression and token reduction make large models lightweight enough for edge-cloud use; agent-layer components (multi-agent frameworks, planning, memory, tool use) turn the model into a service; application-layer batching and applications deliver the user-facing QoS. The paper's discovery is organizational rather than empirical: it claims these bodies of work belong to a single design space and that their integration, not any single technique, is the open research agenda.

Load-bearing premise

The load-bearing premise is that the techniques surveyed at different layers can meaningfully be integrated into one coherent deployment stack for agent services, and that the paper's selection of topics is representative enough to support its conclusions about open problems; the paper does not implement or demonstrate that the layers compose in practice.

Editorial extensions

If this is right

  • If the framework is right, a serving system for FM agents should be designed with cross-layer budgets: a latency or accuracy target at the application layer should be traceable down to choices in execution, resource, and model layers.
  • The survey's own lesson about agent-layer elasticity implies that future serving systems will need adaptive agents that decide when to call APIs, retrieve knowledge, or collaborate with other agents based on current load and task complexity.
  • Edge-cloud deployment of large FMs becomes viable only when parallelism, model compression, and communication optimization are co-designed, since no single hardware class can host the full model.
  • Multi-modal and mixture-of-experts models will require new serving-system mechanisms, because their activated modules and resource demands vary with the input.
  • Batching and scheduling must become heterogeneity-aware, grouping requests by length, service characteristics, and per-request adapters rather than assuming a uniform model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the layered framework implies a concrete design recipe — start from agent-level QoS requirements and derive lower-layer optimization targets — which the paper describes but does not itself validate end-to-end.
  • Editorial inference: the identified agent-layer elasticity gap could be addressed by a scheduler that dynamically selects planning depth, tool-use rate, and collaboration topology under latency constraints; testing such a scheduler would be a natural next step beyond the survey.
  • Editorial inference: cross-layer interactions may create non-compositional effects not quantified in the survey; for example, token reduction changes KV-cache size and attention patterns, which in turn shifts the optimal parallelism and batching strategy.
  • Editorial inference: if the framework is accepted, a useful benchmark would be an open testbed that measures the same agent workload across different layer configurations, making the survey's taxonomy directly actionable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. This survey proposes a five-layer unified framework for deploying foundation model (FM) powered agent services on heterogeneous edge-cloud devices. The layers are: execution optimization (hardware-specific computation, memory and communication optimizations), resource allocation and parallelism (cloud/edge resource scaling and various parallelism strategies), model-layer optimizations (model compression, quantization, distillation, token reduction), agent-layer components (multi-agent frameworks, planning, memory, tool use), and application-layer concerns (batching and representative applications). The paper reviews a large number of recent works in each area, provides several summary tables, and concludes with lessons learned and future research directions. Its central claim is that it is the first comprehensive survey of the deployment of real-time FM-powered agent services in heterogeneous devices.

Significance. The paper has clear value as a broad compilation of recent work on efficient FM inference and on AI-agent architectures. The taxonomy is readable and the tables (e.g., Table II on integrated frameworks, Table XIII on batching) provide useful entry points for readers. The paper is also honest in Section VII about current gaps, which is a strength. However, the significance of the paper as claimed depends entirely on whether the surveyed literature actually constitutes a coherent area of 'FM-powered agent service deployment.' The evidence in the manuscript does not support this: the systems in Sections II-IV are generic FM/LLM inference systems, and the agent work in Section V is algorithmic rather than deployment-oriented. The proposed framework is a plausible future integration agenda, but the claim of a first comprehensive survey of an existing research area is overstated. If the scope were reframed accordingly, the survey would still be a useful reference.

major comments (3)
  1. [I (Introduction, firstness claim) and Sections II-V] The paper's central claim that it is 'the first comprehensive survey to review and discuss the deployment of real-time FM-powered agent services in heterogeneous devices' (Introduction) is not supported by the material surveyed. None of the systems cited in Sections II-IV (e.g., FlashAttention, vLLM, PowerInfer, SpotServe, the resource-allocation works in Table III) is designed for or evaluated on agent workloads such as tool calling, planning loops, multi-agent coordination, or memory retrieval. Conversely, the agent frameworks in Section V (AgentVerse, Toolformer, DEPS, etc.) are described from an algorithmic/application perspective with no system-level deployment contributions. The paper itself concedes in Section VII-A3 that there is a 'significant gap in elasticity at the agent layer' and in Section VII-A1 that heterogeneous edge-cloud FM serving is 'under-explored.' Thus, on the evidence provided, the framework in Figure 1 is a proposal for future integration rather than a taxonomy of an existing body of work on agent-service deployment. This missing evidence is load-bearing because the firstness claim is the paper's principal contribution.
  2. [II-D and III] The connection between the surveyed infrastructure and agent services is never established. For example, Table II lists llama.cpp, MLC-LLM, FastChat, and similar frameworks, but the discussion does not explain how these frameworks support agent-specific requirements such as maintaining multi-turn tool-call state, sharing KV cache across planning iterations, or dynamically deciding when to offload subtasks to different devices. Likewise, the resource-allocation methods in Table III target generic DNN/LLM inference and do not model agent-specific request graphs, inter-agent communication, or memory-retrieval latency. The survey would need at least one worked example or a dedicated analysis showing how the layers compose for an agent service; without this, the unified framework is only a juxtaposition of two adjacent literatures.
  3. [VII-B3 and VII-C] Section VII-B3 lists 'Specific serving system for agents' as a future direction, acknowledging that current serving systems are designed for FM inference rather than for agent services. This is an honest statement, but it directly contradicts the introductory claim that the paper surveys the deployment of FM-powered agent services. The conclusion in Section VII-C repeats that the framework 'showcases the latest advancements' in this area, which is not supported by the content. The authors should either reframe the paper as a survey of building blocks plus a research agenda, or substantially expand the survey to include systems (if any exist) that actually address agent-service deployment end to end.
minor comments (3)
  1. [Throughout] There are many typographical errors and inconsistencies, including stray letters in the author affiliations (e.g., 'Y . Fan', 'V'), inconsistent use of backslashes in Table V, and a duplicated 'Section V' label in Figure 2 (the application layer should be Section VI).
  2. [I] The statistic on ChatGPT users is cited to a non-academic blog-style source (Nerdynav). A more authoritative source, such as a company report or a peer-reviewed citation, would be preferable for a survey.
  3. [V-C] The sentence beginning 'This rethinking method helps...' is grammatically incomplete, and the reference to Hatalis is cited without a first author name or paper title. Please ensure all citations are complete and the prose is polished.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the survey organizes external literature; its firstness claim is a novelty assertion, not a derivation.

full rationale

This is a survey paper, not a derivation or empirical study. It contains no equations that are fitted to data, no parameters calibrated to a subset and then validated on a closely related quantity, and no predictive claims that reduce to its inputs by construction. The central claim is that the paper is the first comprehensive survey of deploying real-time FM-powered agent services in heterogeneous devices; this is a novelty and scope assertion about the literature, not a mathematical or statistical result derived from the framework. The unified framework in Figure 1 is an organizing taxonomy that maps existing work into execution, resource, model, agent, and application layers; such a taxonomy is a presentation device rather than a load-bearing inference. Sections II-IV review external systems for FM inference, compression, parallelism, and resource allocation, while Section V reviews external agent frameworks; none of these sections derives its conclusions from the survey's own framework. Self-citations, such as the prior edge-cloud AIGC survey [7], are used only to position the paper against previous surveys and do not carry the technical weight of the review; even if an author overlaps, the reviewed content consists of independently published systems and is externally checkable. The paper's own limitations, including the stated 'significant gap in elasticity at the agent layer' (Section VII-A3) and the admission that heterogeneous edge computing for FMs is 'under-explored' (Section VII-A1), honestly narrow the claimed scope, which may affect the strength of the firstness claim but is not a form of circularity. No passage defines the survey's key terms in terms of its conclusions, and no cited prior result is invoked as an unexamined theorem to force a choice. The result is therefore self-contained as a literature survey, and the correct circularity finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

As a survey, the paper introduces no new free parameters, axioms, or invented entities. Its central claim is about the existing literature and the usefulness of a proposed organizational framework, which is not the kind of claim that requires a formal axiom ledger.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deploying Foundation Model Powered Agent Services: A Survey." pith.science (2026). https://pith.science/paper/62L5C7UK

@misc{pith2026241213437,
  author       = {Pith},
  title        = {Pith review of: Deploying Foundation Model Powered Agent Services: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/62L5C7UK}},
  note         = {Machine review of arXiv:2412.13437}
}
read the original abstract

Foundation model (FM) powered agent services are regarded as a promising solution to develop intelligent and personalized applications for advancing toward Artificial General Intelligence (AGI). To achieve high reliability and scalability in deploying these agent services, it is essential to collaboratively optimize computational and communication resources, thereby ensuring effective resource allocation and seamless service delivery. In pursuit of this vision, this paper proposes a unified framework aimed at providing a comprehensive survey on deploying FM-based agent services across heterogeneous devices, with the emphasis on the integration of model and resource optimization to establish a robust infrastructure for these services. Particularly, this paper begins with exploring various low-level optimization strategies during inference and studies approaches that enhance system scalability, such as parallelism techniques and resource scaling methods. The paper then discusses several prominent FMs and investigates research efforts focused on inference acceleration, including techniques such as model compression and token reduction. Moreover, the paper also investigates critical components for constructing agent services and highlights notable intelligent applications. Finally, the paper presents potential research directions for developing real-time agent services with high Quality of Service (QoS).

Figures

Figures reproduced from arXiv: 2412.13437 by the authors.

Figure 1
Figure 1. The framework of FM-powered agent services. The execution layer [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The framework of our survey. Each technical session corresponds to a layer in Figure 1. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. A multi-layer optimization framework for edge computing systems [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: The illustration of resource allocation. Resource allocation in a serving [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: The illustration of different parallelism methods. Data parallelism [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 7
Figure 7. Figure 7: The timeline of some popular LLMs. environment [119]. They develop two algorithms: one based on game theory for offline optimization and another leveraging proximal policy optimization for online, adaptive decision￾making processes in a distributed environment. IV. FOU…
Figure 8
Figure 8. Figure 8: The illustration of model adaptation methods. Model selection dy [PITH_FULL_IMAGE:figures/full_fig_p021_8.png]
Figure 9
Figure 9. Figure 9: The illustration of token reduction methods. Token pruning removes [PITH_FULL_IMAGE:figures/full_fig_p022_9.png]
Figure 10
Figure 10. Figure 10: LLM Agent framework. The LLM serves as the central nervous [PITH_FULL_IMAGE:figures/full_fig_p026_10.png]
Figure 11
Figure 11. Figure 11: LLM for planning. Intelligent agents enhance task-handling capa [PITH_FULL_IMAGE:figures/full_fig_p027_11.png]
Figure 12
Figure 12. Figure 12: LLM Agent of Memory. The memory in the historical sequence is [PITH_FULL_IMAGE:figures/full_fig_p028_12.png]
Figure 13
Figure 13. Figure 13: LLM agent of using tools. External APIs and tools can extend the [PITH_FULL_IMAGE:figures/full_fig_p029_13.png]
Figure 14
Figure 14. Figure 14: The illustration of batching. Requests with similar lengths are [PITH_FULL_IMAGE:figures/full_fig_p030_14.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables

    cs.CL 2025-02 conditional novelty 4.0 of 10

    AL4RAG uses a retrieval-aware similarity metric to select annotation-worthy RAG conversation records, yielding DPO-trained models that reject hallucination-prone queries and preserve answer quality.

Reference graph

Works this paper leans on

300 extracted references · 3 canonical work pages · cited by 1 Pith paper

  1. [1]

    On the Opportunities and Risks of Foundation Models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolf- sson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Dur- mus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K...

  2. [2]

    The Rise and Potential of Large Language Model Based Agents: A Survey,

    Z. Xi, W. Chen, X. Guo, W. He, Y . Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, R. Zheng, X. Fan, X. Wang, L. Xiong, Y . Zhou, W. Wang, C. Jiang, Y . Zou, X. Liu, Z. Yin, S. Dou, R. Weng, W. Cheng, Q. Zhang, W. Qin, Y . Zheng, X. Qiu, X. Huang, and T. Gui, “The Rise and Potential of Large Language Model Based Agents: A Survey,” 2023

  3. [3]

    107 Up-to-Date ChatGPT Statistics & User Numbers,

    Nerdynav, “107 Up-to-Date ChatGPT Statistics & User Numbers,” 2024, accessed: 2024-04-24. [Online]. Available: https://nerdynav. com/chatgpt-statistics/

  4. [4]

    A Survey on Hardware Accelerators for Large Language Models,

    C. Kachris, “A Survey on Hardware Accelerators for Large Language Models,” 2024

  5. [5]

    A Survey on Scheduling Techniques in Computing and Network Convergence,

    S. Tang, Y . Yu, H. Wang, G. Wang, W. Chen, Z. Xu, S. Guo, and W. Gao, “A Survey on Scheduling Techniques in Computing and Network Convergence,” IEEE Communications Surveys & Tutorials , vol. 26, no. 1, pp. 160–195, 2024

  6. [6]

    Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems,

    X. Miao, G. Oliaro, Z. Zhang, X. Cheng, H. Jin, T. Chen, and Z. Jia, “Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems,” 2023

  7. [7]

    Unleashing the Power of Edge-Cloud Generative AI in Mobile Net- works: A Survey of AIGC Services,

    M. Xu, H. Du, D. Niyato, J. Kang, Z. Xiong, S. Mao, Z. Han, A. Jamalipour, D. I. Kim, X. Shen, V . C. M. Leung, and H. V . Poor, “Unleashing the Power of Edge-Cloud Generative AI in Mobile Net- works: A Survey of AIGC Services,” IEEE Communications Surveys & Tutorials, pp. 1–1, 2024

  8. [8]

    Machine and Deep Learning for Resource Allocation in Multi-Access Edge Computing: A Survey,

    H. Djigal, J. Xu, L. Liu, and Y . Zhang, “Machine and Deep Learning for Resource Allocation in Multi-Access Edge Computing: A Survey,” IEEE Communications Surveys & Tutorials , vol. 24, no. 4, pp. 2449– 2494, 2022

Show all 300 references
  1. [9]

    A Survey of Large Language Models,

    W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong, Y . Du, C. Yang, Y . Chen, Z. Chen, J. Jiang, R. Ren, Y . Li, X. Tang, Z. Liu, P. Liu, J.-Y . Nie, and J.-R. Wen, “A Survey of Large Language Models,” 2023

  2. [10]

    A Comprehensive Overview of Large Language Models,

    H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian, “A Comprehensive Overview of Large Language Models,” 2024

  3. [11]

    A Survey on Model Compression for Large Language Models,

    X. Zhu, J. Li, Y . Liu, C. Ma, and W. Wang, “A Survey on Model Compression for Large Language Models,” 2023

  4. [12]

    Model Compression and Efficient Inference for Large Language Models: A Survey,

    W. Wang, W. Chen, Y . Luo, Y . Long, Z. Lin, L. Zhang, B. Lin, D. Cai, and X. He, “Model Compression and Efficient Inference for Large Language Models: A Survey,” 2024

  5. [13]

    A survey on knowledge distillation of large language models,

    X. Xu, M. Li, C. Tao, T. Shen, R. Cheng, J. Li, C. Xu, D. Tao, and T. Zhou, “A survey on knowledge distillation of large language models,” arXiv preprint arXiv:2402.13116 , 2024

  6. [14]

    A survey on large language model based autonomous agents,

    L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y . Lin et al. , “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, pp. 1–26, 2024

  7. [15]

    Large language model based multi-agents: A survey of progress and challenges,

    T. Guo, X. Chen, Y . Wang, R. Chang, S. Pei, N. V . Chawla, O. Wiest, and X. Zhang, “Large language model based multi-agents: A survey of progress and challenges,” arXiv preprint arXiv:2402.01680 , 2024

  8. [16]

    FedDSE: Distribution-aware Sub-model Extraction for Fed- erated Learning over Resource-constrained Devices,

    H. Wang, Y . Jia, M. Zhang, Q. Hu, H. Ren, P. Sun, Y . Wen, and T. Zhang, “FedDSE: Distribution-aware Sub-model Extraction for Fed- erated Learning over Resource-constrained Devices,” in Proceedings of the ACM on Web Conference 2024 , 2024, pp. 2902–2913

  9. [17]

    Hardware accelerator for multi-head attention and position-wise feed-forward in the trans- former,

    S. Lu, M. Wang, S. Liang, J. Lin, and Z. Wang, “Hardware accelerator for multi-head attention and position-wise feed-forward in the trans- former,” in 2020 IEEE 33rd International System-on-Chip Conference (SOCC). IEEE, 2020, pp. 84–89

  10. [18]

    Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,

    H. Jang, J. Kim, J.-E. Jo, J. Lee, and J. Kim, “Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,” in Proceedings of the 46th International Symposium on Computer Architecture, 2019, pp. 250–263

  11. [19]

    Npe: An fpga-based overlay processor for natural language processing,

    H. Khan, A. Khan, Z. Khan, L. B. Huang, K. Wang, and L. He, “Npe: An fpga-based overlay processor for natural language processing,” arXiv preprint arXiv:2104.06535 , 2021

  12. [20]

    Dfx: A low-latency multi-fpga appliance for accelerating transformer- based text generation,

    S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y . Kim, “Dfx: A low-latency multi-fpga appliance for accelerating transformer- based text generation,” in 2022 55th IEEE/ACM International Sympo- sium on Microarchitecture (MICRO) . IEEE, 2022, pp. 616–630

  13. [21]

    Transformer- opu: An fpga-based overlay processor for transformer networks,

    Y . Bai, H. Zhou, K. Zhao, J. Chen, J. Yu, and K. Wang, “Transformer- opu: An fpga-based overlay processor for transformer networks,” in 2023 IEEE 31st Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM) . IEEE, 2023, pp. 221–221

  14. [22]

    A cost-efficient fpga imple- mentation of tiny transformer model using neural ode,

    I. Okubo, K. Sugiura, and H. Matsutani, “A cost-efficient fpga imple- mentation of tiny transformer model using neural ode,” arXiv preprint arXiv:2401.02721, 2024

  15. [23]

    Flightllm: Efficient large language model inference with a complete mapping flow on fpga,

    S. Zeng, J. Liu, G. Dai, X. Yang, T. Fu, H. Wang, W. Ma, H. Sun, S. Li, Z. Huang et al. , “Flightllm: Efficient large language model inference with a complete mapping flow on fpga,” arXiv preprint arXiv:2401.03868, 2024

  16. [24]

    Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,

    T. J. Ham, S. J. Jung, S. Kim, Y . H. Oh, Y . Park, Y . Song, J.-H. Park, S. Lee, K. Park, J. W. Lee et al. , “Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA)....

  17. [25]

    Elsa: Hardware-software co-design for efficient, lightweight self- attention mechanism in neural networks,

    T. J. Ham, Y . Lee, S. H. Seo, S. Kim, H. Choi, S. J. Jung, and J. W. Lee, “Elsa: Hardware-software co-design for efficient, lightweight self- attention mechanism in neural networks,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) . IEEE, ...

  18. [26]

    Spatten: Efficient sparse attention architecture with cascade token and head pruning,

    H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2021, pp. 97–110

  19. [27]

    Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,

    L. Lu, Y . Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y . Liang, “Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture, 2021, pp. 977–991

  20. [28]

    Energon: Toward efficient acceleration of transformers using dynamic sparse attention,

    Z. Zhou, J. Liu, Z. Gu, and G. Sun, “Energon: Toward efficient acceleration of transformers using dynamic sparse attention,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 42, no. 1, pp. 136–149, 2022

  21. [29]

    Att: A fault-tolerant reram accelerator for attention-based neural networks,

    H. Guo, L. Peng, J. Zhang, Q. Chen, and T. D. LeCompte, “Att: A fault-tolerant reram accelerator for attention-based neural networks,” in 2020 IEEE 38th International Conference on Computer Design (ICCD). IEEE, 2020, pp. 213–221

  22. [30]

    In-memory com- puting based accelerator for transformer networks for long sequences,

    A. F. Laguna, A. Kazemi, M. Niemier, and X. S. Hu, “In-memory com- puting based accelerator for transformer networks for long sequences,” in 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 2021, pp. 1839–1844

  23. [31]

    Work in progress: Real-time transformer inference on edge ai accelerators,

    B. Reidy, M. Mohammadi, M. Elbtity, H. Smith, and Z. Ramtin, “Work in progress: Real-time transformer inference on edge ai accelerators,” in 2023 IEEE 29th Real-Time and Embedded Technology and Appli- cations Symposium (RTAS) , 2023, pp. 341–344

  24. [32]

    Simplifying transformer blocks,

    B. He and T. Hofmann, “Simplifying transformer blocks,” arXiv preprint arXiv:2311.01906, 2023

  25. [33]

    Accelerating transformer networks through recomposing softmax layers,

    J. Choi, H. Li, B. Kim, S. Hwang, and J. H. Ahn, “Accelerating transformer networks through recomposing softmax layers,” in 2022 IEEE International Symposium on Workload Characterization (IISWC). IEEE, 2022, pp. 92–103

  26. [34]

    Inference with reference: Lossless acceleration of large language models,

    N. Yang, T. Ge, L. Wang, B. Jiao, D. Jiang, L. Yang, R. Majumder, and F. Wei, “Inference with reference: Lossless acceleration of large language models,” arXiv preprint arXiv:2304.04487 , 2023

  27. [35]

    Exponentially faster language mod- elling,

    P. Belcak and R. Wattenhofer, “Exponentially faster language mod- elling,” arXiv preprint arXiv:2311.10770 , 2023

  28. [36]

    Efficient llm inference on cpus,

    H. Shen, H. Chang, B. Dong, Y . Luo, and H. Meng, “Efficient llm inference on cpus,” arXiv preprint arXiv:2311.00502 , 2023

  29. [37]

    Powerinfer: Fast large language model serving with a consumer-grade gpu,

    Y . Song, Z. Mi, H. Xie, and H. Chen, “Powerinfer: Fast large language model serving with a consumer-grade gpu,” arXiv preprint arXiv:2312.12456, 2023

  30. [38]

    Hetegen: Het- erogeneous parallel inference for large language models on resource- constrained devices,

    X. Zhao, B. Jia, H. Zhou, Z. Liu, S. Cheng, and Y . You, “Hetegen: Het- erogeneous parallel inference for large language models on resource- constrained devices,” arXiv preprint arXiv:2403.01164 , 2024

  31. [39]

    Flexgen: High-throughput generative inference of large language models with a single gpu,

    Y . Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Re, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single gpu,” 2023

  32. [40]

    Deja vu: Contextual sparsity MANUSCRIPT 35 for efficient llms at inference time,

    Z. Liu, J. Wang, T. Dao, T. Zhou, B. Yuan, Z. Song, A. Shrivastava, C. Zhang, Y . Tian, C. Re et al. , “Deja vu: Contextual sparsity MANUSCRIPT 35 for efficient llms at inference time,” in International Conference on Machine Learning. PMLR, 2023, pp. 22 137–22 176

  33. [41]

    Llm in a flash: Efficient large language model inference with limited memory,

    K. Alizadeh, I. Mirzadeh, D. Belenko, K. Khatamifard, M. Cho, C. C. Del Mundo, M. Rastegari, and M. Farajtabar, “Llm in a flash: Efficient large language model inference with limited memory,” arXiv preprint arXiv:2312.11514, 2023

  34. [42]

    Fast transformer decoding: One write-head is all you need,

    N. Shazeer, “Fast transformer decoding: One write-head is all you need,” arXiv preprint arXiv:1911.02150 , 2019

  35. [43]

    Gqa: Training generalized multi-query transformer models from multi-head checkpoints,

    J. Ainslie, J. Lee-Thorp, M. de Jong, Y . Zemlyanskiy, F. Lebron, and S. Sanghai, “Gqa: Training generalized multi-query transformer models from multi-head checkpoints,” in The 2023 Conference on Empirical Methods in Natural Language Processing , 2023

  36. [44]

    Efficient memory management for large language model serving with pagedattention,

    W. Kwon, Z. Li, S. Zhuang, Y . Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626

  37. [45]

    Flashattention: Fast and memory-efficient exact attention with io-awareness,

    T. Dao, D. Fu, S. Ermon, A. Rudra, and C. R ´e, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022

  38. [46]

    Flashattention-2: Faster attention with better parallelism and work partitioning,

    T. Dao, “Flashattention-2: Faster attention with better parallelism and work partitioning,” arXiv preprint arXiv:2307.08691 , 2023

  39. [47]

    Flashdecoding++: Faster large language model inference on gpus,

    K. Hong, G. Dai, J. Xu, Q. Mao, X. Li, J. Liu, K. Chen, H. Dong, and Y . Wang, “Flashdecoding++: Faster large language model inference on gpus,” arXiv preprint arXiv:2311.01282 , 2023

  40. [48]

    Bminf: An efficient toolkit for big model inference and tuning,

    X. Han, G. Zeng, W. Zhao, Z. Liu, Z. Zhang, J. Zhou, J. Zhang, J. Chao, and M. Sun, “Bminf: An efficient toolkit for big model inference and tuning,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, 2022, pp. 224– 230

  41. [49]

    Splitwise: Efficient generative llm inference using phase splitting,

    P. Patel, E. Choukse, C. Zhang, ´I. Goiri, A. Shah, S. Maleki, and R. Bianchini, “Splitwise: Efficient generative llm inference using phase splitting,” arXiv preprint arXiv:2311.18677 , 2023

  42. [50]

    Fast distributed inference serving for large language models,

    B. Wu, Y . Zhong, Z. Zhang, G. Huang, X. Liu, and X. Jin, “Fast distributed inference serving for large language models,” arXiv preprint arXiv:2305.05920, 2023

  43. [51]

    Specinfer: Accelerating generative large language model serving with speculative inference and token tree verification,

    X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, R. Y . Y . Wong, A. Zhu, L. Yang, X. Shi, C. Shi, Z. Chen, D. Arfeen, R. Abhyankar, and Z. Jia, “Specinfer: Accelerating generative large language model serving with speculative inference and token tree verification,” 2023

  44. [52]

    Llmcad: Fast and scalable on-device large language model inference,

    D. Xu, W. Yin, X. Jin, Y . Zhang, S. Wei, M. Xu, and X. Liu, “Llmcad: Fast and scalable on-device large language model inference,” arXiv preprint arXiv:2309.04255, 2023

  45. [53]

    Deepspeed-moe: Advancing mixture- of-experts inference and training to power next-generation ai scale,

    S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y . Aminabadi, A. A. Awan, J. Rasley, and Y . He, “Deepspeed-moe: Advancing mixture- of-experts inference and training to power next-generation ai scale,” in International conference on machine learning . PMLR, 2022, pp. 18 332–18 346

  46. [54]

    Edgemoe: Fast on-device inference of moe-based large language models,

    R. Yi, L. Guo, S. Wei, A. Zhou, S. Wang, and M. Xu, “Edgemoe: Fast on-device inference of moe-based large language models,” arXiv preprint arXiv:2308.14352, 2023

  47. [55]

    Fast inference of mixture-of-experts lan- guage models with offloading,

    A. Eliseev and D. Mazur, “Fast inference of mixture-of-experts lan- guage models with offloading,”arXiv preprint arXiv:2312.17238, 2023

  48. [56]

    Moe-infinity: Activation- aware expert offloading for efficient moe serving,

    L. Xue, Y . Fu, Z. Lu, L. Mai, and M. Marina, “Moe-infinity: Activation- aware expert offloading for efficient moe serving,” arXiv preprint arXiv:2401.14361, 2024

  49. [57]

    Fiddler: Cpu-gpu orchestration for fast inference of mixture-of-experts models,

    K. Kamahori, Y . Gu, K. Zhu, and B. Kasikci, “Fiddler: Cpu-gpu orchestration for fast inference of mixture-of-experts models,” arXiv preprint arXiv:2402.07033, 2024

  50. [58]

    Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference,

    R. Hwang, J. Wei, S. Cao, C. Hwang, X. Tang, T. Cao, M. Yang, and M. Rhu, “Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference,” arXiv preprint arXiv:2308.12066, 2023

  51. [59]

    Intelligence-endogenous management platform for computing and network convergence,

    Z. Hong, X. Qiu, J. Lin, W. Chen, Y . Yu, H. Wang, S. Guo, and W. Gao, “Intelligence-endogenous management platform for computing and network convergence,” IEEE Network, 2023

  52. [60]

    Resource allocation in large language model integrated 6g vehicular networks,

    C. Liu and J. Zhao, “Resource allocation in large language model integrated 6g vehicular networks,” arXiv preprint arXiv:2403.19016 , 2024

  53. [61]

    Lingualinked: A distributed large language model inference system for mobile de- vices,

    J. Zhao, Y . Song, S. Liu, I. G. Harris, and S. A. Jyothi, “Lingualinked: A distributed large language model inference system for mobile de- vices,” arXiv preprint arXiv:2312.00388 , 2023

  54. [62]

    {MegaScale}: Scaling large language model training to more than 10,000 {GPUs},

    Z. Jiang, H. Lin, Y . Zhong, Q. Huang, Y . Chen, Z. Zhang, Y . Peng, X. Li, C. Xie, S. Nong et al. , “ {MegaScale}: Scaling large language model training to more than 10,000 {GPUs},” in 21st USENIX Sym- posium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp...

  55. [63]

    LOSP: Overlap synchronization parallel with local compensation for fast dis- tributed training,

    H. Wang, Z. Qu, S. Guo, N. Wang, R. Li, and W. Zhuang, “LOSP: Overlap synchronization parallel with local compensation for fast dis- tributed training,” IEEE Journal on Selected Areas in Communications, vol. 39, no. 8, pp. 2541–2557, 2021

  56. [64]

    Heterogeneous semantic and bit communications: A semi-noma scheme,

    X. Mu, Y . Liu, L. Guo, and N. Al-Dhahir, “Heterogeneous semantic and bit communications: A semi-noma scheme,” IEEE Journal on Selected Areas in Communications , vol. 41, no. 1, pp. 155–169, 2022

  57. [65]

    Computing networks enabled semantic communications,

    Z. Qin, J. Ying, D. Yang, H. Wang, and X. Tao, “Computing networks enabled semantic communications,” IEEE Network, 2024

  58. [66]

    ggerganov/llama.cpp: Port of facebook’s llama model in c/c++

    G. Gerganov, “ggerganov/llama.cpp: Port of facebook’s llama model in c/c++.” https://github.com/ggerganov/llama.cpp, 2023

  59. [67]

    MLC-LLM,

    M. team, “MLC-LLM,” 2023. [Online]. Available: https://github.com/ mlc-ai/mlc-llm

  60. [68]

    mnn-llm: llm deploy project based mnn

    mnn llm, “mnn-llm: llm deploy project based mnn.” https://github.com/ wangzhaode/mnn-llm, 2023

  61. [69]

    Judging llm-as-a-judge with mt-bench and chatbot arena,

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023

  62. [70]

    Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

    R. Y . Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al., “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Com- ...

  63. [71]

    Openvino deep learning workbench: Comprehensive analysis and tuning of neural networks inference,

    Y . Gorbachev, M. Fedorov, I. Slavutin, A. Tugarev, M. Fatekhov, and Y . Tarkan, “Openvino deep learning workbench: Comprehensive analysis and tuning of neural networks inference,” inProceedings of the IEEE/CVF International Conference on Computer Vision Workshops , 2019, pp. 0–0

  64. [72]

    mllm is a fast and lightweight multimodal llm inference engine for mobile and edge devices

    mllm, “mllm is a fast and lightweight multimodal llm inference engine for mobile and edge devices.” https://github.com/UbiquitousLearning/ mllm, 2023

  65. [73]

    Fp6-llm: Efficiently serving large language models through fp6-centric algorithm-system co-design,

    H. Xia, Z. Zheng, X. Wu, S. Chen, Z. Yao, S. Youn, A. Bakhtiari, M. Wyatt, D. Zhuang, Z. Zhouet al., “Fp6-llm: Efficiently serving large language models through fp6-centric algorithm-system co-design,” arXiv preprint arXiv:2401.14112 , 2024

  66. [74]

    Colossal-ai: A unified deep learning system for large-scale parallel training,

    S. Li, H. Liu, Z. Bian, J. Fang, H. Huang, Y . Liu, B. Wang, and Y . You, “Colossal-ai: A unified deep learning system for large-scale parallel training,” in Proceedings of the 52nd International Conference on Parallel Processing, 2023, pp. 766–775

  67. [75]

    Efficient large-scale language model training on gpu clusters using megatron-lm,

    D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V . Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro et al. , “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference fo...

  68. [76]

    A tensorrt toolbox for optimized large language model inference,

    tensorrtllm, “A tensorrt toolbox for optimized large language model inference,” https://github.com/NVIDIA/TensorRT-LLM, 2023

  69. [77]

    Langchain,

    Harrison Chase, “Langchain,” https://github.com/langchain-ai/ langchain, 2024, accessed: 2024-04-07

  70. [78]

    Parrot: Efficient Serving of LLM-based Applications with Semantic Variable,

    C. Lin, Z. Han, C. Zhang, Y . Yang, F. Yang, C. Chen, and L. Qiu, “Parrot: Efficient Serving of LLM-based Applications with Semantic Variable,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara, CA: USENIX Associa- tion, Jul. 2024

  71. [79]

    SGLang: Efficient Execution of Structured Language Model Pro- grams,

    L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y . Sheng, “SGLang: Efficient Execution of Structured Language Model Pro- grams,” 2024

  72. [80]

    Clipper: A {Low-Latency} online prediction serving system,

    D. Crankshaw, X. Wang, G. Zhou, M. J. Franklin, J. E. Gonzalez, and I. Stoica, “Clipper: A {Low-Latency} online prediction serving system,” in 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) , 2017, pp. 613–627

  73. [81]

    {MArk}: Exploiting cloud services for {Cost-Effective},{SLO-Aware} machine learning infer- ence serving,

    C. Zhang, M. Yu, W. Wang, and F. Yan, “ {MArk}: Exploiting cloud services for {Cost-Effective},{SLO-Aware} machine learning infer- ence serving,” in 2019 USENIX Annual Technical Conference (USENIX ATC 19), 2019, pp. 1049–1062

  74. [82]

    Nexus: A GPU cluster engine for accelerating DNN-based video analysis,

    H. Shen, L. Chen, Y . Jin, L. Zhao, B. Kong, M. Philipose, A. Kr- ishnamurthy, and R. Sundaram, “Nexus: A GPU cluster engine for accelerating DNN-based video analysis,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles, 2019, pp. 322–337

  75. [83]

    InferLine: latency-aware provisioning and scaling for prediction serving pipelines,

    D. Crankshaw, G.-E. Sela, X. Mo, C. Zumar, I. Stoica, J. Gonzalez, and A. Tumanov, “InferLine: latency-aware provisioning and scaling for prediction serving pipelines,” in Proceedings of the 11th ACM Symposium on Cloud Computing , 2020, pp. 477–491

  76. [84]

    Serving {DNNs} like clockwork: Performance predictability from the bottom up,

    A. Gujarati, R. Karimi, S. Alzayat, W. Hao, A. Kaufmann, Y . Vig- fusson, and J. Mace, “Serving {DNNs} like clockwork: Performance predictability from the bottom up,” in 14th USENIX Symposium on MANUSCRIPT 36 Operating Systems Design and Implementation (OSDI 20) , 2020, pp. 443–462

  77. [85]

    {INFaaS}: Automated model-less inference serving,

    F. Romero, Q. Li, N. J. Yadwadkar, and C. Kozyrakis, “ {INFaaS}: Automated model-less inference serving,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21) , 2021, pp. 397–411

  78. [86]

    Morphling: Fast, near-optimal auto-configuration for cloud- native model serving,

    L. Wang, L. Yang, Y . Yu, W. Wang, B. Li, X. Sun, J. He, and L. Zhang, “Morphling: Fast, near-optimal auto-configuration for cloud- native model serving,” inProceedings of the ACM Symposium on Cloud Computing, 2021, pp. 639–653

  79. [87]

    Cocktail: A multidimensional optimization for model serving in cloud,

    Gunasekaran, Jashwant Raj and Mishra, Cyan Subhra and Thinakaran, Prashanth and Sharma, Bikash and Kandemir, Mahmut Taylan and Das, Chita R, “Cocktail: A multidimensional optimization for model serving in cloud,” in 19th USENIX Symposium on Networked Systems Design and Impleme...

  80. [88]

    Kairos: Building cost- efficient machine learning inference systems with heterogeneous cloud resources,

    B. Li, S. Samsi, V . Gadepally, and D. Tiwari, “Kairos: Building cost- efficient machine learning inference systems with heterogeneous cloud resources,” in Proceedings of the 32nd International Symposium on High-Performance Parallel and Distributed Computing , 2023, pp. 3– 16

  81. [89]

    {SHEPHERD}: Serving {DNNs} in the wild,

    H. Zhang, Y . Tang, A. Khandelwal, and I. Stoica, “ {SHEPHERD}: Serving {DNNs} in the wild,” in 20th USENIX Symposium on Net- worked Systems Design and Implementation (NSDI 23), 2023, pp. 787– 808

  82. [90]

    Spotserve: Serving generative large language models on preemptible instances,

    X. Miao, C. Shi, J. Duan, X. Xi, D. Lin, B. Cui, and Z. Jia, “Spotserve: Serving generative large language models on preemptible instances,” arXiv preprint arXiv:2311.15566 , 2023

  83. [91]

    Frequency resource allocation and interference management in mobile edge com- puting for an Internet of Things system,

    W. Na, S. Jang, Y . Lee, L. Park, N.-N. Dao, and S. Cho, “Frequency resource allocation and interference management in mobile edge com- puting for an Internet of Things system,” IEEE Internet of Things Journal, vol. 6, no. 3, pp. 4910–4920, 2018

  84. [92]

    Decentralized resource auctioning for latency-sensitive edge computing,

    C. Avasalcai, C. Tsigkanos, and S. Dustdar, “Decentralized resource auctioning for latency-sensitive edge computing,” in 2019 IEEE inter- national conference on edge computing (EDGE) . IEEE, 2019, pp. 72–76

  85. [93]

    Joint computation partitioning and resource allocation for latency sensitive applications in mobile edge clouds,

    L. Yang, B. Liu, J. Cao, Y . Sahni, and Z. Wang, “Joint computation partitioning and resource allocation for latency sensitive applications in mobile edge clouds,” IEEE Transactions on Services Computing , vol. 14, no. 5, pp. 1439–1452, 2019

  86. [94]

    Adaptive computation offloading and resource allocation strategy in a mobile edge computing environment,

    Z. Tong, X. Deng, F. Ye, S. Basodi, X. Xiao, and Y . Pan, “Adaptive computation offloading and resource allocation strategy in a mobile edge computing environment,”Information Sciences, vol. 537, pp. 116– 131, 2020

  87. [95]

    Resource allocation based on deep reinforcement learning in IoT edge computing,

    X. Xiong, K. Zheng, L. Lei, and L. Hou, “Resource allocation based on deep reinforcement learning in IoT edge computing,” IEEE Journal on Selected Areas in Communications , vol. 38, no. 6, pp. 1133–1146, 2020

  88. [96]

    CE-IoT: Cost-effective cloud- edge resource provisioning for heterogeneous IoT applications,

    Z. Zhou, S. Yu, W. Chen, and X. Chen, “CE-IoT: Cost-effective cloud- edge resource provisioning for heterogeneous IoT applications,” IEEE Internet of Things Journal , vol. 7, no. 9, pp. 8600–8614, 2020

  89. [97]

    Dynamic resource allocation and computation offloading for IoT fog computing system,

    Z. Chang, L. Liu, X. Guo, and Q. Sheng, “Dynamic resource allocation and computation offloading for IoT fog computing system,” IEEE Transactions on Industrial Informatics , vol. 17, no. 5, pp. 3348–3357, 2020

  90. [98]

    Lass: Running latency sensitive serverless computations at the edge,

    B. Wang, A. Ali-Eldin, and P. Shenoy, “Lass: Running latency sensitive serverless computations at the edge,” in Proceedings of the 30th international symposium on high-performance parallel and distributed computing, 2021, pp. 239–251

  91. [99]

    Resource provisioning and allocation in function- as-a-service edge-clouds,

    O. Ascigil, A. G. Tasiopoulos, T. K. Phan, V . Sourlas, I. Psaras, and G. Pavlou, “Resource provisioning and allocation in function- as-a-service edge-clouds,” IEEE Transactions on Services Computing , vol. 15, no. 4, pp. 2410–2424, 2021

  92. [100]

    KneeScale: Efficient resource scaling for serverless computing at the edge,

    X. Li, P. Kang, J. Molone, W. Wang, and P. Lama, “KneeScale: Efficient resource scaling for serverless computing at the edge,” in2022 22nd IEEE International Symposium on Cluster, Cloud and Internet Computing (CCGrid). IEEE, 2022, pp. 180–189

  93. [101]

    CEC: A containerized edge computing framework for dynamic resource provisioning,

    S. Hu, W. Shi, and G. Li, “CEC: A containerized edge computing framework for dynamic resource provisioning,” IEEE Transactions on Mobile Computing, 2022

  94. [102]

    FedCDA: Federated Learning with Cross-rounds Divergence-aware Aggregation,

    H. Wang, H. Xu, Y . Li, Y . Xu, R. Li, and T. Zhang, “FedCDA: Federated Learning with Cross-rounds Divergence-aware Aggregation,” in The Twelfth International Conference on Learning Representations , 2024

  95. [103]

    Megatron-lm: Training multi-billion parameter language models using model parallelism,

    M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catan- zaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” arXiv preprint arXiv:1909.08053 , 2019

  96. [104]

    Efficiently scaling transformer inference,

    R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” Proceedings of Machine Learning and Systems , vol. 5, 2023

  97. [105]

    Alpaserve: Statistical multiplexing with model parallelism for deep learning serving,

    Z. Li, L. Zheng, Y . Zhong, V . Liu, Y . Sheng, X. Jin, Y . Huang, Z. Chen, H. Zhang, J. E. Gonzalez et al. , “Alpaserve: Statistical multiplexing with model parallelism for deep learning serving,” in 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2...

  98. [106]

    Lightseq: Sequence level parallelism for distributed training of long context transformers,

    D. Li, R. Shao, A. Xie, E. P. Xing, J. E. Gonzalez, I. Stoica, X. Ma, and H. Zhang, “Lightseq: Sequence level parallelism for distributed training of long context transformers,” arXiv preprint arXiv:2310.03294, 2023

  99. [107]

    Distributed Inference and Fine-tuning of Large Language Models Over The In- ternet,

    A. Borzunov, M. Ryabinin, A. Chumachenko, D. Baranchuk, T. Dettmers, Y . Belkada, P. Samygin, and C. A. Raffel, “Distributed Inference and Fine-tuning of Large Language Models Over The In- ternet,” Advances in Neural Information Processing Systems , vol. 36, 2024

  100. [108]

    Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills,

    A. Agrawal, A. Panwar, J. Mohan, N. Kwatra, B. S. Gulavani, and R. Ramjee, “Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills,” arXiv preprint arXiv:2308.16369 , 2023

  101. [109]

    DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving,

    Y . Zhong, S. Liu, J. Chen, J. Hu, Y . Zhu, X. Liu, X. Jin, and H. Zhang, “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving,” arXiv preprint arXiv:2401.09670, 2024

  102. [110]

    Distributing deep neural networks with containerized partitions at the edge,

    L. Zhou, H. Wen, R. Teodorescu, and D. H. Du, “Distributing deep neural networks with containerized partitions at the edge,” in 2nd USENIX Workshop on Hot Topics in Edge Computing (HotEdge 19) , 2019

  103. [111]

    Distributed inference acceleration with adaptive DNN partitioning and offloading,

    T. Mohammed, C. Joe-Wong, R. Babbar, and M. Di Francesco, “Distributed inference acceleration with adaptive DNN partitioning and offloading,” in IEEE INFOCOM 2020-IEEE Conference on Computer Communications. IEEE, 2020, pp. 854–863

  104. [112]

    Coedge: Cooperative dnn inference with adaptive workload partitioning over heterogeneous edge devices,

    L. Zeng, X. Chen, Z. Zhou, L. Yang, and J. Zhang, “Coedge: Cooperative dnn inference with adaptive workload partitioning over heterogeneous edge devices,” IEEE/ACM Transactions on Networking, vol. 29, no. 2, pp. 595–608, 2020

  105. [113]

    Throughput maximization of delay-aware DNN inference in edge computing by exploring DNN model partitioning and inference parallelism,

    J. Li, W. Liang, Y . Li, Z. Xu, X. Jia, and S. Guo, “Throughput maximization of delay-aware DNN inference in edge computing by exploring DNN model partitioning and inference parallelism,” IEEE Transactions on Mobile Computing , vol. 22, no. 5, pp. 3017–3030, 2021

  106. [114]

    Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures,

    Y . Wu, M. Lentz, D. Zhuo, and Y . Lu, “Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures,” Pro- ceedings of the VLDB Endowment , vol. 16, no. 3, pp. 406–419, 2022

  107. [115]

    PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices,

    Y . Hu, C. Imes, X. Zhao, S. Kundu, P. A. Beerel, S. P. Crago, and J. P. Walters, “PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices,” in 2022 25th Euromicro Conference on Digital System Design (DSD) , 2022, pp. 298–307

  108. [116]

    PDD: partitioning DAG-topology DNNs for streaming tasks,

    L. Wu, G. Gao, J. Yu, F. Zhou, Y . Yang, and T. Wang, “PDD: partitioning DAG-topology DNNs for streaming tasks,” IEEE Internet of Things Journal , 2023

  109. [117]

    Dnn partitioning for inference throughput acceleration at the edge,

    T. Feltin, L. March ´o, J.-A. Cordero-Fuertes, F. Brockners, and T. H. Clausen, “Dnn partitioning for inference throughput acceleration at the edge,” IEEE Access, vol. 11, pp. 52 236–52 249, 2023

  110. [118]

    Distributed DNN Inference with Fine-grained Model Partitioning in Mobile Edge Computing Networks,

    H. Li, X. Li, Q. Fan, Q. He, X. Wang, and V . C. Leung, “Distributed DNN Inference with Fine-grained Model Partitioning in Mobile Edge Computing Networks,” IEEE Transactions on Mobile Computing , 2024

  111. [119]

    MoEI: Mobility-Aware Edge Inference Based on Model Partition and Service Migration,

    Z. Liu, M. Tian, M. Dong, X. Wang, C. Qiu, and C. Zhang, “MoEI: Mobility-Aware Edge Inference Based on Model Partition and Service Migration,” IEEE Transactions on Mobile Computing, no. 01, pp. 1–14, 2024

  112. [120]

    NVIDIA Triton Inference Server,

    NVIDIA Corporation, “NVIDIA Triton Inference Server,” https: //developer.nvidia.com/nvidia-triton-inference-server, 2024, accessed: 2024-04-17

  113. [121]

    TensorFlow Serving,

    Google LLC, “TensorFlow Serving,” https://www.tensorflow.org/tfx/ guide/serving, 2024, accessed: 2024-04-17

  114. [122]

    FedNLR: Federated Learning with Neuron-wise Learning Rates,

    H. Wang, P. Zheng, X. Han, W. Xu, R. Li, and T. Zhang, “FedNLR: Federated Learning with Neuron-wise Learning Rates,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024, pp. 3069–3080

  115. [123]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017

  116. [124]

    Exploring the limits of transfer learning with a unified text-to-text transformer,

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research, vol. 21, no. 140, pp. 1–67, 2020. MANUSCRIPT 37

  117. [125]

    Language models are unsupervised multitask learners

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners.”

  118. [126]

    Language mod- els are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language mod- els are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020

  119. [127]

    Pangu- α: Large-scale autoregressive pretrained chinese language models with auto-parallel computation,

    W. Zeng, X. Ren, T. Su, H. Wang, Y . Liao, Z. Wang, X. Jiang, Z. Yang, K. Wang, X. Zhang et al. , “Pangu- α: Large-scale autoregressive pretrained chinese language models with auto-parallel computation,” arXiv e-prints, pp. arXiv–2104, 2021

  120. [128]

    Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation,

    Y . Sun, S. Wang, S. Feng, S. Ding, C. Pang, J. Shang, J. Liu, X. Chen, Y . Zhao, Y . Lu et al. , “Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation,” arXiv preprint arXiv:2107.02137, 2021

  121. [129]

    Jurassic-1: Technical details and evaluation,

    O. Lieber, O. Sharir, B. Lenz, and Y . Shoham, “Jurassic-1: Technical details and evaluation,” White Paper. AI21 Labs, vol. 1, p. 9, 2021

  122. [130]

    Training compute-optimal large language models,

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clarket al., “Training compute-optimal large language models,” arXiv preprint arXiv:2203.15556, 2022

  123. [131]

    Lamda: Language models for dialog applications,

    R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y . Duet al., “Lamda: Language models for dialog applications,” arXiv preprint arXiv:2201.08239 , 2022

  124. [132]

    Us- ing deepspeed and megatron to train megatron-turing nlg 530b, a large- scale generative language model,

    S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V . Korthikantiet al., “Us- ing deepspeed and megatron to train megatron-turing nlg 530b, a large- scale generative language model,” arXiv preprint arXiv:2201.11990 , 2022

  125. [133]

    Scaling language models: Methods, analysis & insights from training gopher,

    J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young et al., “Scaling language models: Methods, analysis & insights from training gopher,” arXiv preprint arXiv:2112.11446, 2021

  126. [134]

    Opt: Open pre-trained transformer language models,

    S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V . Lin et al. , “Opt: Open pre-trained transformer language models,” arXiv preprint arXiv:2205.01068, 2022

  127. [135]

    Galactica: A large language model for science,

    R. Taylor, M. Kardas, G. Cucurull, T. Scialom, A. Hartshorn, E. Sar- avia, A. Poulton, V . Kerkez, and R. Stojnic, “Galactica: A large language model for science,” arXiv preprint arXiv:2211.09085 , 2022

  128. [136]

    Bloom: A 176b-parameter open-access multilingual language model,

    T. Le Scao, A. Fan, C. Akiki, E. Pavlick, S. Ili ´c, D. Hesslow, R. Castagn ´e, A. S. Luccioni, F. Yvon, M. Gall ´e et al. , “Bloom: A 176b-parameter open-access multilingual language model,” 2023

  129. [137]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023

  130. [138]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al. , “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023

  131. [139]

    Baichuan-7b: About: A large-scale 7b pretrain- ing language model developed by baichuan-inc,

    BaiChuan-Inc, “Baichuan-7b: About: A large-scale 7b pretrain- ing language model developed by baichuan-inc,” https://github.com/ baichuan-inc/Baichuan-7B/tree/main, 2024, accessed: 2024-04-07

  132. [140]

    A 13b large language model developed by baichuan intelligent technology,

    Baichuan Intelligent Technology, “A 13b large language model developed by baichuan intelligent technology,” https://github.com/ baichuan-inc/Baichuan-13B/tree/main, 2024, accessed: 2024-04-07

  133. [141]

    Qwen technical report,

    J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huang et al. , “Qwen technical report,” arXiv preprint arXiv:2309.16609, 2023

  134. [142]

    Skywork: A more open bilingual foundation model,

    T. Wei, L. Zhao, L. Zhang, B. Zhu, L. Wang, H. Yang, B. Li, C. Cheng, W. L ¨u, R. Hu et al. , “Skywork: A more open bilingual foundation model,” arXiv preprint arXiv:2310.19341 , 2023

  135. [143]

    The falcon series of open language models,

    E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojo- caru, M. Debbah, ´E. Goffinet, D. Hesslow, J. Launay, Q. Malartic et al. , “The falcon series of open language models,” arXiv preprint arXiv:2311.16867, 2023

  136. [144]

    Starcoder: may the source be with you!

    R. Li, L. B. Allal, Y . Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim et al., “Starcoder: may the source be with you!” arXiv preprint arXiv:2305.06161 , 2023

  137. [145]

    Yi: Open foundation models by 01. ai,

    A. Young, B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, H. Li, J. Zhu, J. Chen, J. Chang et al., “Yi: Open foundation models by 01. ai,” arXiv preprint arXiv:2403.04652 , 2024

  138. [146]

    Finetuned language models are zero-shot learners,

    J. Wei, M. Bosma, V . Y . Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V . Le, “Finetuned language models are zero-shot learners,” 2022

  139. [147]

    mt5: A massively multilingual pre-trained text-to-text transformer,

    L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,” arXiv preprint arXiv:2010.11934 , 2020

  140. [148]

    Flan-moe: Scaling instruction- finetuned language models with sparse mixture of experts,

    S. Shen, L. Hou, Y . Zhou, N. Du, S. Longpre, J. Wei, H. W. Chung, B. Zoph, W. Fedus, X. Chen et al. , “Flan-moe: Scaling instruction- finetuned language models with sparse mixture of experts,” arXiv e- prints, pp. arXiv–2305, 2023

  141. [149]

    Scaling instruction- finetuned language models,

    H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y . Tay, W. Fedus, Y . Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V . Zhao, Y . Huang, A. Dai, H. Y...

  142. [150]

    Taori, I

    R. Taori, I. Gulrajani, T. Zhang, Y . Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Alpaca,” 2023. [Online]. Available: https://crfm.stanford.edu/2023/03/13/alpaca.html

  143. [151]

    Glm-130b: An open bilingual pre- trained model,

    A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y . Xu, W. Zheng, X. Xia, W. L. Tam, Z. Ma, Y . Xue, J. Zhai, W. Chen, P. Zhang, Y . Dong, and J. Tang, “Glm-130b: An open bilingual pre- trained model,” 2023

  144. [152]

    Flm-101b: An open llm and how to train it with $100 k budget,

    X. Li, Y . Yao, X. Jiang, X. Fang, X. Meng, S. Fan, P. Han, J. Li, L. Du, B. Qin et al., “Flm-101b: An open llm and how to train it with $100 k budget,” arXiv preprint arXiv:2309.03852 , 2023

  145. [153]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems , vol. 35, pp. 23 716–23 736, 2022

  146. [154]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023

  147. [155]

    Minigpt-4: En- hancing vision-language understanding with advanced large language models,

    D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: En- hancing vision-language understanding with advanced large language models,” arXiv preprint arXiv:2304.10592 , 2023

  148. [156]

    mplug-owl: Modularization empowers large language models with multimodality,

    Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y . Zhou, J. Wang, A. Hu, P. Shi, Y . Shi et al. , “mplug-owl: Modularization empowers large language models with multimodality,” arXiv preprint arXiv:2304.14178 , 2023

  149. [157]

    Cogvlm: Visual expert for pretrained language models,

    W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y . Wang, J. Ji, Z. Yang, L. Zhao, X. Song et al., “Cogvlm: Visual expert for pretrained language models,” arXiv preprint arXiv:2311.03079 , 2023

  150. [158]

    Imagebind: One embedding space to bind them all,

    R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 15 180–15 190

  151. [159]

    Pandagpt: One model to instruction-follow them all,

    Y . Su, T. Lan, H. Li, J. Xu, Y . Wang, and D. Cai, “Pandagpt: One model to instruction-follow them all,” arXiv preprint arXiv:2305.16355, 2023

  152. [160]

    Next-gpt: Any-to-any multimodal llm,

    S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua, “Next-gpt: Any-to-any multimodal llm,” arXiv preprint arXiv:2309.05519 , 2023

  153. [161]

    Onellm: One framework to align all modalities with language,

    J. Han, K. Gong, Y . Zhang, J. Wang, K. Zhang, D. Lin, Y . Qiao, P. Gao, and X. Yue, “Onellm: One framework to align all modalities with language,” arXiv preprint arXiv:2312.03700 , 2023

  154. [162]

    Gemini: a family of highly capable multimodal models,

    G. Team, R. Anil, S. Borgeaud, Y . Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth et al., “Gemini: a family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805 , 2023

  155. [163]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,

    M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser et al. , “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” arXiv preprint arXiv:2403.05530 , 2024

  156. [164]

    Shikra: Unleashing multimodal llm’s referential dialogue magic,

    K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao, “Shikra: Unleashing multimodal llm’s referential dialogue magic,” arXiv preprint arXiv:2306.15195 , 2023

  157. [165]

    Ferret: Refer and ground anything anywhere at any granularity,

    H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S.-F. Chang, and Y . Yang, “Ferret: Refer and ground anything anywhere at any granularity,” arXiv preprint arXiv:2310.07704 , 2023

  158. [166]

    Pix2seq: A language modeling framework for object detection,

    T. Chen, S. Saxena, L. Li, D. J. Fleet, and G. Hinton, “Pix2seq: A language modeling framework for object detection,” 2022

  159. [167]

    Next- chat: An lmm for chat, detection and segmentation,

    A. Zhang, L. Zhao, C.-W. Xie, Y . Zheng, W. Ji, and T.-S. Chua, “Next- chat: An lmm for chat, detection and segmentation,” arXiv preprint arXiv:2311.04498, 2023

  160. [168]

    Pangu- {\Sigma}: To- wards trillion parameter language model with sparse heterogeneous computing,

    X. Ren, P. Zhou, X. Meng, X. Huang, Y . Wang, W. Wang, P. Li, X. Zhang, A. Podolskiy, G. Arshinov et al. , “Pangu- {\Sigma}: To- wards trillion parameter language model with sparse heterogeneous computing,” arXiv preprint arXiv:2303.10845 , 2023

  161. [169]

    Mistral 7b,

    A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier et al., “Mistral 7b,” arXiv preprint arXiv:2310.06825 , 2023

  162. [170]

    Mixtral of experts,

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bam- ford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand et al. , “Mixtral of experts,” arXiv preprint arXiv:2401.04088 , 2024. MANUSCRIPT 38

  163. [171]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

    W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, 2022

  164. [172]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” 2019

  165. [173]

    Tinybert: Distilling bert for natural language understanding,

    X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, “Tinybert: Distilling bert for natural language understanding,” arXiv preprint arXiv:1909.10351 , 2019

  166. [174]

    Rethinking optimization and architecture for tiny language models,

    Y . Tang, F. Liu, Y . Ni, Y . Tian, Z. Bai, Y .-Q. Hu, S. Liu, S. Jui, K. Han, and Y . Wang, “Rethinking optimization and architecture for tiny language models,” arXiv preprint arXiv:2402.02791 , 2024

  167. [175]

    Tinyllama: An open-source small language model,

    P. Zhang, G. Zeng, T. Wang, and W. Lu, “Tinyllama: An open-source small language model,” arXiv preprint arXiv:2401.02385 , 2024

  168. [176]

    Phi- 3 technical report: A highly capable language model locally on your phone,

    M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl et al. , “Phi- 3 technical report: A highly capable language model locally on your phone,” arXiv preprint arXiv:2404.14219 , 2024

  169. [177]

    A simple and effec- tive pruning approach for large language models,

    M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effec- tive pruning approach for large language models,” arXiv preprint arXiv:2306.11695, 2023

  170. [178]

    Llm-pruner: On the structural pruning of large language models,

    X. Ma, G. Fang, and X. Wang, “Llm-pruner: On the structural pruning of large language models,” Advances in neural information processing systems, vol. 36, 2024

  171. [179]

    Loraprune: Pruning meets low-rank parameter-efficient fine-tuning,

    M. Zhang, H. Chen, C. Shen, Z. Yang, L. Ou, X. Yu, and B. Zhuang, “Loraprune: Pruning meets low-rank parameter-efficient fine-tuning,” 2023

  172. [180]

    Lorashear: Efficient large language model structured pruning and knowledge recovery,

    T. Chen, T. Ding, B. Yadav, I. Zharkov, and L. Liang, “Lorashear: Efficient large language model structured pruning and knowledge recovery,” arXiv preprint arXiv:2310.18356 , 2023

  173. [181]

    Fluctuation-based adaptive structured pruning for large language models,

    Y . An, X. Zhao, T. Yu, M. Tang, and J. Wang, “Fluctuation-based adaptive structured pruning for large language models,” arXiv preprint arXiv:2312.11983, 2023

  174. [182]

    One-shot sensitivity-aware mixed sparsity pruning for large language models,

    H. Shao, B. Liu, and Y . Qian, “One-shot sensitivity-aware mixed sparsity pruning for large language models,” arXiv preprint arXiv:2310.09499, 2023

  175. [183]

    Compresso: Structured pruning with collaborative prompting learns compact large language models,

    S. Guo, J. Xu, L. L. Zhang, and M. Yang, “Compresso: Structured pruning with collaborative prompting learns compact large language models,” arXiv preprint arXiv:2310.05015 , 2023

  176. [184]

    Sheared llama: Accelerating language model pre-training via structured pruning,

    M. Xia, T. Gao, Z. Zeng, and D. Chen, “Sheared llama: Accelerating language model pre-training via structured pruning,” arXiv preprint arXiv:2310.06694, 2023

  177. [185]

    Pruning large language models via accuracy predictor,

    Y . Ji, Y . Cao, and J. Liu, “Pruning large language models via accuracy predictor,” arXiv preprint arXiv:2309.09507 , 2023

  178. [186]

    Beyond size: How gradients shape pruning decisions in large language models,

    R. J. Das, L. Ma, and Z. Shen, “Beyond size: How gradients shape pruning decisions in large language models,” arXiv preprint arXiv:2311.04902, 2023

  179. [187]

    Dynamic context pruning for efficient and interpretable autoregressive transformers,

    S. Anagnostidis, D. Pavllo, L. Biggio, L. Noci, A. Lucchi, and T. Hofmann, “Dynamic context pruning for efficient and interpretable autoregressive transformers,” Advances in Neural Information Process- ing Systems, vol. 36, 2024

  180. [188]

    ZipLM: Inference-Aware Struc- tured Pruning of Language Models,

    E. Kurti ´c, E. Frantar, and D. Alistarh, “ZipLM: Inference-Aware Struc- tured Pruning of Language Models,” Advances in Neural Information Processing Systems, vol. 36, 2024

  181. [189]

    DPHuBERT: Joint distillation and pruning of self-supervised speech models,

    Y . Peng, Y . Sudo, S. Muhammad, and S. Watanabe, “DPHuBERT: Joint distillation and pruning of self-supervised speech models,” arXiv preprint arXiv:2305.17651, 2023

  182. [190]

    Structured Prun- ing of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding,

    Y . Peng, K. Kim, F. Wu, P. Sridhar, and S. Watanabe, “Structured Prun- ing of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5

  183. [191]

    MoPE-CLIP: Structured Pruning for Efficient Vision- Language Models with Module-wise Pruning Error Metric,

    H. Lin, H. Bai, Z. Liu, L. Hou, M. Sun, L. Song, Y . Wei, and Z. Sun, “MoPE-CLIP: Structured Pruning for Efficient Vision- Language Models with Module-wise Pruning Error Metric,” arXiv preprint arXiv:2403.07839, 2024

  184. [192]

    X-pruner: explainable pruning for vision trans- formers,

    L. Yu and W. Xiang, “X-pruner: explainable pruning for vision trans- formers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 24 355–24 363

  185. [193]

    Structural pruning for diffusion models,

    G. Fang, X. Ma, and X. Wang, “Structural pruning for diffusion models,” Advances in neural information processing systems , vol. 36, 2024

  186. [194]

    A unified pruning framework for vision transform- ers,

    H. Yu and J. Wu, “A unified pruning framework for vision transform- ers,” Science China Information Sciences , vol. 66, no. 7, p. 179101, 2023

  187. [195]

    CAP: Correlation-Aware Pruning for Highly-Accurate Sparse Vision Mod- els,

    D. Kuznedelev, E. Kurti ´c, E. Frantar, and D. Alistarh, “CAP: Correlation-Aware Pruning for Highly-Accurate Sparse Vision Mod- els,” in Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curr...

  188. [196]

    Pada: Pruning assisted domain adaptation for self-supervised speech representations,

    V . S. Lodagala, S. Ghosh, and S. Umesh, “Pada: Pruning assisted domain adaptation for self-supervised speech representations,” in 2022 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2023, pp. 136–143

  189. [197]

    Smoothquant: Accurate and efficient post-training quantization for large language models,

    G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning. PMLR, 2023, pp. 38 087–38 099

  190. [198]

    Rptq: Reorder-based post-training quantization for large language models,

    Z. Yuan, L. Niu, J. Liu, W. Liu, X. Wang, Y . Shang, G. Sun, Q. Wu, J. Wu, and B. Wu, “Rptq: Reorder-based post-training quantization for large language models,” arXiv preprint arXiv:2304.01089 , 2023

  191. [199]

    Loftq: Lora-fine-tuning-aware quantization for large language models,

    Y . Li, Y . Yu, C. Liang, P. He, N. Karampatziakis, W. Chen, and T. Zhao, “Loftq: Lora-fine-tuning-aware quantization for large language models,” arXiv preprint arXiv:2310.08659 , 2023

  192. [200]

    Outlier suppression+: Accurate quantization of large language mod- els by equivalent and optimal shifting and scaling,

    X. Wei, Y . Zhang, Y . Li, X. Zhang, R. Gong, J. Guo, and X. Liu, “Outlier suppression+: Accurate quantization of large language mod- els by equivalent and optimal shifting and scaling,” arXiv preprint arXiv:2304.09145, 2023

  193. [201]

    FPTQ: Fine-grained Post-Training Quantization for Large Language Models,

    Q. Li, Y . Zhang, L. Li, P. Yao, B. Zhang, X. Chu, Y . Sun, L. Du, and Y . Xie, “FPTQ: Fine-grained Post-Training Quantization for Large Language Models,” arXiv preprint arXiv:2308.15987 , 2023

  194. [202]

    Owq: Lessons learned from activation outliers for weight quantization in large language models,

    C. Lee, J. Jin, T. Kim, H. Kim, and E. Park, “Owq: Lessons learned from activation outliers for weight quantization in large language models,” arXiv preprint arXiv:2306.02272 , 2023

  195. [203]

    Awq: Activation-aware weight quantization for llm compression and accel- eration,

    J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, and S. Han, “Awq: Activation-aware weight quantization for llm compression and accel- eration,” arXiv preprint arXiv:2306.00978 , 2023

  196. [204]

    Integer or floating point? new outlooks for low-bit quantization on large language models,

    Y . Zhang, L. Zhao, S. Cao, W. Wang, T. Cao, F. Yang, M. Yang, S. Zhang, and N. Xu, “Integer or floating point? new outlooks for low-bit quantization on large language models,” arXiv preprint arXiv:2305.12356, 2023

  197. [205]

    Omniquant: Omnidirectionally calibrated quan- tization for large language models,

    W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, K. Zhang, P. Gao, Y . Qiao, and P. Luo, “Omniquant: Omnidirectionally calibrated quan- tization for large language models,” arXiv preprint arXiv:2308.13137 , 2023

  198. [206]

    IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact,

    R. Liu, H. Bai, H. Lin, Y . Li, H. Gao, Z. Xu, L. Hou, J. Yao, and C. Yuan, “IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact,” arXiv preprint arXiv:2403.01241, 2024

  199. [207]

    Memory-Efficient Fine-Tuning of Compressed Large Language Mod- els via sub-4-bit Integer Quantization,

    J. Kim, J. H. Lee, S. Kim, J. Park, K. M. Yoo, S. J. Kwon, and D. Lee, “Memory-Efficient Fine-Tuning of Compressed Large Language Mod- els via sub-4-bit Integer Quantization,” in Advances in Neural Informa- tion Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M...

  200. [208]

    Qllm: Accurate and efficient low-bitwidth quantization for large language models,

    J. Liu, R. Gong, X. Wei, Z. Dong, J. Cai, and B. Zhuang, “Qllm: Accurate and efficient low-bitwidth quantization for large language models,” arXiv preprint arXiv:2310.08041 , 2023

  201. [209]

    Llm-qat: Data-free quantization aware training for large language models,

    Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y . Mehdad, Y . Shi, R. Kr- ishnamoorthi, and V . Chandra, “Llm-qat: Data-free quantization aware training for large language models,” arXiv preprint arXiv:2305.17888, 2023

  202. [210]

    Quip: 2-bit quanti- zation of large language models with guarantees,

    J. Chee, Y . Cai, V . Kuleshov, and C. M. De Sa, “Quip: 2-bit quanti- zation of large language models with guarantees,” Advances in Neural Information Processing Systems , vol. 36, 2024

  203. [211]

    Norm tweaking: High- performance low-bit quantization of large language models,

    L. Li, Q. Li, B. Zhang, and X. Chu, “Norm tweaking: High- performance low-bit quantization of large language models,” arXiv preprint arXiv:2309.02784, 2023

  204. [212]

    Zeroquant-v2: Exploring post-training quantization in llms from comprehensive study to low rank compensation,

    Z. Yao, X. Wu, C. Li, S. Youn, and Y . He, “Zeroquant-v2: Exploring post-training quantization in llms from comprehensive study to low rank compensation,” arXiv preprint arXiv:2303.08302 , 2023

  205. [213]

    Qa-lora: Quantization-aware low-rank adaptation of large language models,

    Y . Xu, L. Xie, X. Gu, X. Chen, H. Chang, H. Zhang, Z. Chen, X. Zhang, and Q. Tian, “Qa-lora: Quantization-aware low-rank adaptation of large language models,” arXiv preprint arXiv:2309.14717 , 2023

  206. [214]

    Int2.1: Towards fine-tunable quantized large language models with error cor- rection through low-rank adaptation,

    Y . Chai, J. Gkountouras, G. G. Ko, D. Brooks, and G.-Y . Wei, “Int2.1: Towards fine-tunable quantized large language models with error cor- rection through low-rank adaptation,”arXiv preprint arXiv:2306.08162, 2023

  207. [215]

    The Quantization Model of Neural Scaling,

    E. J. Michaud, Z. Liu, U. Girit, and M. Tegmark, “The Quantization Model of Neural Scaling,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023

  208. [216]

    Quantized transformer language model implementations on edge devices,

    M. W. U. Rahman, M. M. Abrar, H. G. Copening, S. Hariri, S. Shao, P. Satam, and S. Salehi, “Quantized transformer language model implementations on edge devices,” arXiv preprint arXiv:2310.03971 , 2023. MANUSCRIPT 39

  209. [217]

    Distilling step-by-step! outper- forming larger language models with less training data and smaller model sizes,

    C.-Y . Hsieh, C.-L. Li, C.-K. Yeh, H. Nakhost, Y . Fujii, A. Ratner, R. Krishna, C.-Y . Lee, and T. Pfister, “Distilling step-by-step! outper- forming larger language models with less training data and smaller model sizes,” arXiv preprint arXiv:2305.02301 , 2023

  210. [218]

    Zephyr: Direct distillation of lm alignment,

    L. Tunstall, E. Beeching, N. Lambert, N. Rajani, K. Rasul, Y . Belkada, S. Huang, L. von Werra, C. Fourrier, N. Habib et al., “Zephyr: Direct distillation of lm alignment,” arXiv preprint arXiv:2310.16944 , 2023

  211. [219]

    Lion: Adversarial distillation of closed-source large language model,

    Y . Jiang, C. Chan, M. Chen, and W. Wang, “Lion: Adversarial distillation of closed-source large language model,” arXiv preprint arXiv:2305.12870, 2023

  212. [220]

    Pad: Program- aided distillation specializes large models in reasoning,

    X. Zhu, B. Qi, K. Zhang, X. Long, and B. Zhou, “Pad: Program- aided distillation specializes large models in reasoning,” arXiv preprint arXiv:2305.13888, 2023

  213. [221]

    Distillspec: Improving speculative decoding via knowledge distillation,

    Y . Zhou, K. Lyu, A. S. Rawat, A. K. Menon, A. Rostamizadeh, S. Ku- mar, J.-F. Kagy, and R. Agarwal, “Distillspec: Improving speculative decoding via knowledge distillation,” arXiv preprint arXiv:2310.08461, 2023

  214. [222]

    Knowledge distillation of llm for education,

    E. Latif, L. Fang, P. Ma, and X. Zhai, “Knowledge distillation of llm for education,” arXiv preprint arXiv:2312.15842 , 2023

  215. [223]

    Less is more: Task-aware layer-wise distillation for language model compres- sion,

    C. Liang, S. Zuo, Q. Zhang, P. He, W. Chen, and T. Zhao, “Less is more: Task-aware layer-wise distillation for language model compres- sion,” in International Conference on Machine Learning . PMLR, 2023, pp. 20 852–20 867

  216. [224]

    Homodistil: Homotopic task-agnostic distillation of pre-trained transformers,

    C. Liang, H. Jiang, Z. Li, X. Tang, B. Yin, and T. Zhao, “Homodistil: Homotopic task-agnostic distillation of pre-trained transformers,” arXiv preprint arXiv:2302.09632, 2023

  217. [225]

    Symbolic chain-of-thought distillation: Small models can also

    L. H. Li, J. Hessel, Y . Yu, X. Ren, K.-W. Chang, and Y . Choi, “Symbolic chain-of-thought distillation: Small models can also” think” step-by-step,” arXiv preprint arXiv:2306.14050 , 2023

  218. [226]

    Evolving Knowledge Distillation with Large Language Models and Active Learning,

    C. Liu, Y . Kang, F. Zhao, K. Kuang, Z. Jiang, C. Sun, and F. Wu, “Evolving Knowledge Distillation with Large Language Models and Active Learning,” arXiv preprint arXiv:2403.06414 , 2024

  219. [227]

    Minimal Distillation Schedule for Extreme Language Model Com- pression,

    C. Zhang, Y . Yang, Q. Wang, J. Liu, J. Wang, W. Wu, and D. Song, “Minimal Distillation Schedule for Extreme Language Model Com- pression,” in Findings of the Association for Computational Linguistics: EACL 2024, 2024, pp. 1378–1394

  220. [228]

    Slam: Student-label mixing for distillation with unlabeled ex- amples,

    V . Kontonis, F. Iliopoulos, K. Trinh, C. Baykal, G. Menghani, and E. Vee, “Slam: Student-label mixing for distillation with unlabeled ex- amples,” Advances in Neural Information Processing Systems , vol. 36, 2024

  221. [229]

    Universalner: Targeted distillation from large language models for open named entity recognition,

    W. Zhou, S. Zhang, Y . Gu, M. Chen, and H. Poon, “Universalner: Targeted distillation from large language models for open named entity recognition,” arXiv preprint arXiv:2308.03279 , 2023

  222. [230]

    Scott: Self-consistent chain-of-thought distillation,

    P. Wang, Z. Wang, Z. Li, Y . Gao, B. Yin, and X. Ren, “Scott: Self-consistent chain-of-thought distillation,” arXiv preprint arXiv:2305.01879, 2023

  223. [231]

    Detrdistill: A universal knowledge distillation framework for detr- families,

    J. Chang, S. Wang, H.-M. Xu, Z. Chen, C. Yang, and F. Zhao, “Detrdistill: A universal knowledge distillation framework for detr- families,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 6898–6908

  224. [232]

    Boosting accuracy and robustness of student models via adaptive adversarial distillation,

    B. Huang, M. Chen, Y . Wang, J. Lu, M. Cheng, and W. Wang, “Boosting accuracy and robustness of student models via adaptive adversarial distillation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 24 668–24 677

  225. [233]

    Dime-fm: Distilling multimodal and efficient foundation models,

    X. Sun, P. Zhang, P. Zhang, H. Shah, K. Saenko, and X. Xia, “Dime-fm: Distilling multimodal and efficient foundation models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 15 521–15 533

  226. [234]

    Contextualization distillation from large language model for knowledge graph completion,

    D. Li, Z. Tan, T. Chen, and H. Liu, “Contextualization distillation from large language model for knowledge graph completion,” arXiv preprint arXiv:2402.01729, 2024

  227. [235]

    Large Language Model Meets Graph Neural Network in Knowledge Distillation,

    S. Hu, G. Zou, S. Yang, B. Zhang, and Y . Chen, “Large Language Model Meets Graph Neural Network in Knowledge Distillation,” arXiv preprint arXiv:2402.05894, 2024

  228. [236]

    On Good Practices for Task-Specific Distillation of Large Pretrained Models,

    J. Marrie, M. Arbel, J. Mairal, and D. Larlus, “On Good Practices for Task-Specific Distillation of Large Pretrained Models,” arXiv preprint arXiv:2402.11305, 2024

  229. [237]

    Dafkd: Domain-aware federated knowledge distillation,

    H. Wang, Y . Li, W. Xu, R. Li, Y . Zhan, and Z. Zeng, “Dafkd: Domain-aware federated knowledge distillation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2023, pp. 20 412–20 421

  230. [238]

    Edgeadaptor: Online configuration adaption, model selection and resource provisioning for edge dnn inference serving at scale,

    K. Zhao, Z. Zhou, X. Chen, R. Zhou, X. Zhang, S. Yu, and D. Wu, “Edgeadaptor: Online configuration adaption, model selection and resource provisioning for edge dnn inference serving at scale,” IEEE Transactions on Mobile Computing , 2022

  231. [239]

    Sti: Turbocharge nlp inference at the edge via elastic pipelining,

    L. Guo, W. Choe, and F. X. Lin, “Sti: Turbocharge nlp inference at the edge via elastic pipelining,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2023, pp. 791–803

  232. [240]

    Tabi: An efficient multi- level inference system for large language models,

    Y . Wang, K. Chen, H. Tan, and K. Guo, “Tabi: An efficient multi- level inference system for large language models,” in Proceedings of the Eighteenth European Conference on Computer Systems , 2023, pp. 233–248

  233. [241]

    Consistent Acceler- ated Inference via Confident Adaptive Transformers,

    T. Schuster, A. Fisch, T. Jaakkola, and R. Barzilay, “Consistent Acceler- ated Inference via Confident Adaptive Transformers,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 4962–4979

  234. [242]

    Virtual reality solutions employing artificial intelligence methods: A systematic literature review,

    T. Ribeiro de Oliveira, B. Biancardi Rodrigues, M. Moura da Silva, R. Antonio N. Spinass ´e, G. Giesen Ludke, M. Ruy Soares Gaudio, G. Iglesias Rocha Gomes, L. Guio Cotini, D. da Silva Vargens, M. Queiroz Schimidt et al. , “Virtual reality solutions employing artificial intell...

  235. [243]

    Use of AI-based tools for healthcare purposes: a survey study from consumers’ perspectives,

    P. Esmaeilzadeh, “Use of AI-based tools for healthcare purposes: a survey study from consumers’ perspectives,” BMC medical informatics and decision making , vol. 20, pp. 1–19, 2020

  236. [244]

    Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding,

    Z. Chen, A. May, R. Svirschevski, Y . Huang, M. Ryabinin, Z. Jia, and B. Chen, “Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding,” arXiv preprint arXiv:2402.12374 , 2024

  237. [245]

    Minions: Accelerating Large Language Model Inference with Adaptive and Collective Speculative Decoding,

    S. Wang, H. Yang, X. Wang, T. Liu, P. Wang, X. Liang, K. Ma, T. Feng, X. You, Y . Baoet al., “Minions: Accelerating Large Language Model Inference with Adaptive and Collective Speculative Decoding,” arXiv preprint arXiv:2402.15678, 2024

  238. [246]

    Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs,

    S. Ge, Y . Zhang, L. Liu, M. Zhang, J. Han, and J. Gao, “Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs,” in The Twelfth International Conference on Learning Representations , 2023

  239. [247]

    Efficient streaming language models with attention sinks,

    G. Xiao, Y . Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in The Twelfth International Conference on Learning Representations , 2023

  240. [248]

    H2o: Heavy-hitter oracle for efficient generative inference of large language models,

    Z. Zhang, Y . Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y . Tian, C. R´e, C. Barrett et al., “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024

  241. [249]

    Constraint-aware and ranking-distilled token pruning for efficient transformer inference,

    J. Li, L. L. Zhang, J. Xu, Y . Wang, S. Yan, Y . Xia, Y . Yang, T. Cao, H. Sun, W. Deng et al., “Constraint-aware and ranking-distilled token pruning for efficient transformer inference,” in Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining ,...

  242. [250]

    LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models,

    H. Jiang, Q. Wu, C.-Y . Lin, Y . Yang, and L. Qiu, “LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 13 358–13 376

  243. [251]

    Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression,

    H. Jiang, Q. Wu, X. Luo, D. Li, C.-Y . Lin, Y . Yang, and L. Qiu, “Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression,” arXiv preprint arXiv:2310.06839 , 2023

  244. [252]

    Compressing Context to En- hance Inference Efficiency of Large Language Models,

    Y . Li, B. Dong, F. Guerin, and C. Lin, “Compressing Context to En- hance Inference Efficiency of Large Language Models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 6342–6353

  245. [253]

    Learned token pruning for transformers,

    S. Kim, S. Shen, D. Thorsley, A. Gholami, W. Kwon, J. Hassoun, and K. Keutzer, “Learned token pruning for transformers,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022, pp. 784–794

  246. [254]

    AdapLeR: Speeding up Inference by Adaptive Length Reduction,

    A. Modarressi, H. Mohebbi, and M. T. Pilehvar, “AdapLeR: Speeding up Inference by Adaptive Length Reduction,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1–15

  247. [255]

    Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search,

    G. Kim and K. Cho, “Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume...

  248. [256]

    Power-bert: Accelerating bert inference via progressive word-vector elimination,

    S. Goyal, A. R. Choudhury, S. Raje, V . Chakaravarthy, Y . Sabharwal, and A. Verma, “Power-bert: Accelerating bert inference via progressive word-vector elimination,” in International Conference on Machine Learning. PMLR, 2020, pp. 3690–3699

  249. [257]

    In-context Autoencoder for Context Compression in a Large Language Model,

    T. Ge, H. Jing, L. Wang, X. Wang, S.-Q. Chen, and F. Wei, “In-context Autoencoder for Context Compression in a Large Language Model,” in The Twelfth International Conference on Learning Representations , 2023. MANUSCRIPT 40

  250. [258]

    Learning to compress prompts with gist tokens,

    J. Mu, X. Li, and N. Goodman, “Learning to compress prompts with gist tokens,” Advances in Neural Information Processing Systems , vol. 36, 2024

  251. [259]

    Adapting Language Models to Compress Contexts,

    A. Chevalier, A. Wettig, A. Ajith, and D. Chen, “Adapting Language Models to Compress Contexts,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 3829–3846

  252. [260]

    SYNERGISTIC PATCH PRUN- ING FOR VISION TRANS-FORMER: UNIFYING INTRA-&INTER- LAYER PATCH IMPORTANCE

    Y . Zhang, L. Wei, and N. M. Freris, “SYNERGISTIC PATCH PRUN- ING FOR VISION TRANS-FORMER: UNIFYING INTRA-&INTER- LAYER PATCH IMPORTANCE.”

  253. [261]

    A Simple Romance Between Multi-Exit Vision Transformer and Token Reduction,

    D. Liu, M. Kan, S. Shan, and C. Xilin, “A Simple Romance Between Multi-Exit Vision Transformer and Token Reduction,” in The Twelfth International Conference on Learning Representations , 2023

  254. [262]

    Heatvit: Hardware-efficient adaptive token prun- ing for vision transformers,

    P. Dong, M. Sun, A. Lu, Y . Xie, K. Liu, Z. Kong, X. Meng, Z. Li, X. Lin, Z. Fang et al., “Heatvit: Hardware-efficient adaptive token prun- ing for vision transformers,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2023, pp. 442–455

  255. [263]

    Patch slimming for efficient vision transformers,

    Y . Tang, K. Han, Y . Wang, C. Xu, J. Guo, C. Xu, and D. Tao, “Patch slimming for efficient vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 165–12 174

  256. [264]

    Dynamicvit: Efficient vision transformers with dynamic token sparsification,

    Y . Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “Dynamicvit: Efficient vision transformers with dynamic token sparsification,” Ad- vances in neural information processing systems , vol. 34, pp. 13 937– 13 949, 2021

  257. [265]

    Token Merging: Your ViT But Faster,

    D. Bolya, C.-Y . Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token Merging: Your ViT But Faster,” in The Eleventh International Conference on Learning Representations , 2022

  258. [266]

    EViT: Expediting Vision Transformers via Token Reorganizations,

    Y . Liang, G. Chongjian, Z. Tong, Y . Song, J. Wang, and P. Xie, “EViT: Expediting Vision Transformers via Token Reorganizations,” in International Conference on Learning Representations , 2021

  259. [267]

    DiffRate: Differentiable Compression Rate for Efficient Vision Transformers,

    M. Chen, W. Shao, P. Xu, M. Lin, K. Zhang, F. Chao, R. Ji, Y . Qiao, and P. Luo, “DiffRate: Differentiable Compression Rate for Efficient Vision Transformers,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, 2023, pp. 17 118–17 128

  260. [268]

    Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision Transformers,

    S. Long, Z. Zhao, J. Pi, S. Wang, and J. Wang, “Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision Transformers,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2023, pp. 10 334– 10 343

  261. [269]

    Joint token pruning and squeezing towards more aggressive compression of vision trans- formers,

    S. Wei, T. Ye, S. Zhang, Y . Tang, and J. Liang, “Joint token pruning and squeezing towards more aggressive compression of vision trans- formers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2092–2101

  262. [270]

    PuMer: Pruning and Merging Tokens for Efficient Vision Language Models,

    Q. Cao, B. Paranjape, and H. Hajishirzi, “PuMer: Pruning and Merging Tokens for Efficient Vision Language Models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12 890–12 903

  263. [271]

    TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding,

    S. Ren, S. Chen, S. Li, X. Sun, and L. Hou, “TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp. 932–947

  264. [272]

    Prune spatio-temporal tokens by semantic-aware temporal accumulation,

    S. Ding, P. Zhao, X. Zhang, R. Qian, H. Xiong, and Q. Tian, “Prune spatio-temporal tokens by semantic-aware temporal accumulation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 16 945–16 956

  265. [273]

    Efficient video transformers with spatial-temporal token selection,

    J. Wang, X. Yang, H. Li, L. Liu, Z. Wu, and Y .-G. Jiang, “Efficient video transformers with spatial-temporal token selection,” in European Conference on Computer Vision . Springer, 2022, pp. 69–86

  266. [274]

    Efficient Transform- ers: A Survey,

    Y . Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient Transform- ers: A Survey,” ACM Comput. Surv., vol. 55, no. 6, dec 2022

  267. [275]

    Which tokens to use? investigating token reduction in vision transformers,

    J. B. Haurum, S. Escalera, G. W. Taylor, and T. B. Moeslund, “Which tokens to use? investigating token reduction in vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 773–783

  268. [276]

    OTAS: An Elastic Transformer Serving System via Token Adapta- tion,

    J. Chen, W. Xu, Z. Hong, S. Guo, H. Wang, J. Zhang, and D. Zeng, “OTAS: An Elastic Transformer Serving System via Token Adapta- tion,” 2024

  269. [277]

    Adaptive computation with elastic input sequence,

    F. Xue, V . Likhosherstov, A. Arnab, N. Houlsby, M. Dehghani, and Y . You, “Adaptive computation with elastic input sequence,” in Inter- national Conference on Machine Learning. PMLR, 2023, pp. 38 971– 38 988

  270. [278]

    Levels of AGI: Operationalizing Progress on the Path to AGI,

    M. R. Morris, J. Sohl-dickstein, N. Fiedel, T. Warkentin, A. Dafoe, A. Faust, C. Farabet, and S. Legg, “Levels of AGI: Operationalizing Progress on the Path to AGI,” arXiv preprint arXiv:2311.02462, 2023

  271. [279]

    AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors,

    W. Chen, Y . Su, J. Zuo, C. Yang, C. Yuan, C.-M. Chan, yang Yu, Y . Lu, Y .-H. Hung, C. Qian, Y . Qin, X. Cong, R. Xie, Z. Liu, M. Sun, and J. Zhou, “AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors,” in The Twelfth International Conference o...

  272. [280]

    Agentbench: Evaluating llms as agents,

    X. Liu, H. Yu, H. Zhang, Y . Xu, X. Lei, H. Lai, Y . Gu, H. Ding, K. Men, K. Yang et al., “Agentbench: Evaluating llms as agents,” arXiv preprint arXiv:2308.03688, 2023

  273. [281]

    Building cooperative embodied agents modularly with large language models,

    H. Zhang, W. Du, J. Shan, Q. Zhou, Y . Du, J. B. Tenenbaum, T. Shu, and C. Gan, “Building cooperative embodied agents modularly with large language models,” arXiv preprint arXiv:2307.02485 , 2023

  274. [282]

    Evaluating multi-agent coordination abilities in large language models,

    S. Agashe, Y . Fan, and X. E. Wang, “Evaluating multi-agent coordination abilities in large language models,” arXiv preprint arXiv:2310.03903, 2023

  275. [283]

    Grounded Decod- ing: Guiding Text Generation with Grounded Models for Embodied Agents,

    W. Huang, F. Xia, D. Shah, D. Driess, A. Zeng, Y . Lu, P. Florence, I. Mordatch, S. Levine, K. Hausman, and B. Ichter, “Grounded Decod- ing: Guiding Text Generation with Grounded Models for Embodied Agents,” in Conference on Neural Information Processing Systems , 2023

  276. [284]

    Empowering Conversational Agents using Semantic In-Context Learning,

    A. Omidvar and A. An, “Empowering Conversational Agents using Semantic In-Context Learning,” in Annual Meeting of the Association for Computational Linguistics , 2023

  277. [285]

    Multi-agent collaboration: Harnessing the power of intelligent llm agents,

    Y . Talebirad and A. Nadiri, “Multi-agent collaboration: Harnessing the power of intelligent llm agents,” arXiv preprint arXiv:2306.03314, 2023

  278. [286]

    Dynamic llm-agent network: An llm-agent collaboration framework with agent team opti- mization,

    Z. Liu, Y . Zhang, P. Li, Y . Liu, and D. Yang, “Dynamic llm-agent network: An llm-agent collaboration framework with agent team opti- mization,” arXiv preprint arXiv:2310.02170 , 2023

  279. [287]

    Mindagent: Emergent gaming interaction,

    R. Gong, Q. Huang, X. Ma, H. V o, Z. Durante, Y . Noda, Z. Zheng, S.-C. Zhu, D. Terzopoulos, L. Fei-Fei et al. , “Mindagent: Emergent gaming interaction,” arXiv preprint arXiv:2309.09971 , 2023

  280. [288]

    Describe, explain, plan and select: interactive planning with LLMs enables open- world multi-task agents,

    Z. Wang, S. Cai, G. Chen, A. Liu, X. S. Ma, and Y . Liang, “Describe, explain, plan and select: interactive planning with LLMs enables open- world multi-task agents,” Advances in Neural Information Processing Systems, vol. 36, 2024

  281. [289]

    Large language models as common- sense knowledge for large-scale task planning,

    Z. Zhao, W. S. Lee, and D. Hsu, “Large language models as common- sense knowledge for large-scale task planning,” Advances in Neural Information Processing Systems , vol. 36, 2024

  282. [290]

    Large language models can implement policy iteration,

    E. Brooks, L. Walls, R. L. Lewis, and S. Singh, “Large language models can implement policy iteration,” Advances in Neural Information Processing Systems, vol. 36, 2024

  283. [291]

    Large language models of code fail at completing code with potential bugs,

    T. Dinh, J. Zhao, S. Tan, R. Negrinho, L. Lausen, S. Zha, and G. Karypis, “Large language models of code fail at completing code with potential bugs,” Advances in Neural Information Processing Systems, vol. 36, 2024

  284. [292]

    Lever- aging pre-trained large language models to construct and utilize world models for model-based task planning,

    L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati, “Lever- aging pre-trained large language models to construct and utilize world models for model-based task planning,” Advances in Neural Informa- tion Processing Systems , vol. 36, 2024

  285. [293]

    Avalon’s Game of Thoughts: Battle Against Deception through Recursive Contemplation,

    S. Wang, C. Liu, Z. Zheng, S. Qi, S. Chen, Q. Yang, A. Zhao, C. Wang, S. Song, and G. Huang, “Avalon’s Game of Thoughts: Battle Against Deception through Recursive Contemplation,” arXiv preprint arXiv:2310.01320, 2023

  286. [294]

    Prompted LLMs as Chatbot Modules for Long Open-domain Conversation,

    G. Lee, V . Hartmann, J. Park, D. Papailiopoulos, and K. Lee, “Prompted LLMs as Chatbot Modules for Long Open-domain Conversation,” in Annual Meeting of the Association for Computational Linguistics , 2023

  287. [295]

    Large Lan- guage Models Are Semi-Parametric Reinforcement Learning Agents,

    D. Zhang, L. Chen, S. Zhang, H. Xu, Z. Zhao, and K. Yu, “Large Lan- guage Models Are Semi-Parametric Reinforcement Learning Agents,” Advances in Neural Information Processing Systems , vol. 36, 2024

  288. [296]

    Augmenting language models with long-term memory,

    W. Wang, L. Dong, H. Cheng, X. Liu, X. Yan, J. Gao, and F. Wei, “Augmenting language models with long-term memory,” Advances in Neural Information Processing Systems , vol. 36, 2024

  289. [297]

    Memory-augmented llm personalization with short-and long-term memory coordination,

    K. Zhang, F. Zhao, Y . Kang, and X. Liu, “Memory-augmented llm personalization with short-and long-term memory coordination,” arXiv preprint arXiv:2309.11696, 2023

  290. [298]

    Memory Matters: The Need to Improve Long-Term Memory in LLM-Agents,

    K. Hatalis, D. Christou, J. Myers, S. Jones, K. Lambert, A. Amos- Binks, Z. Dannenhauer, and D. Dannenhauer, “Memory Matters: The Need to Improve Long-Term Memory in LLM-Agents,” in Proceedings of the AAAI Symposium Series , vol. 2, no. 1, 2023, pp. 277–280

  291. [299]

    Mot: Memory-of-thought enables chatgpt to self-improve,

    X. Li and X. Qiu, “Mot: Memory-of-thought enables chatgpt to self-improve,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 6354–6374

  292. [300]

    From llm to conversational agent: A memory enhanced architecture with fine-tuning of large language models,

    N. Liu, L. Chen, X. Tian, W. Zou, K. Chen, and M. Cui, “From llm to conversational agent: A memory enhanced architecture with fine-tuning of large language models,” arXiv preprint arXiv:2401.02777 , 2024. MANUSCRIPT 41

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.