REVIEW 3 major objections 3 minor 1 cited by
Deploying Foundation Model Powered Agent Services: A Survey
T0 review · 3 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This survey argues that delivering real-time foundation-model agent services at scale requires a unified deployment stack linking hardware execution, resource management, model compression, agent components, and applications.
desk verdict A useful survey of LLM serving and agent frameworks whose framing as a survey of 'agent-service deployment' overstates how much the cited work is actually about agents; the paper's own lessons section concedes the integration is open. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the five-layer framework in the paper's Figure 1: an Execution layer, a Resource layer, a Model layer, an Agent layer, and an Application layer. It functions as both a taxonomy and a compositional claim: low-level inference optimizations, resource-allocation and parallelism strategies, model compression and token reduction, agent capabilities, and batching are treated as mutually dependent design choices within one serving stack. The framework carries the argument by showing where each surveyed technique sits and by exposing the missing elasticity at the agent layer as the binding constraint on real-time agent services.
What would settle it
A concrete check would be to search the literature for a prior survey that already covers real-time FM-powered agent deployment across heterogeneous edge-cloud devices under a unified framework; if one exists, the paper's first-comprehensive claim fails. A second check would be an end-to-end experiment combining representative techniques from all five layers (say token reduction, pipeline parallelism, and agent tool calling) to see whether their benefits add up or interfere; the paper reports no such experiment.
Extended reading notes
Core claim
On its own terms, the paper's central claim is that it provides the first comprehensive survey of real-time FM-powered agent service deployment across heterogeneous devices, and that this deployment is best understood through a unified framework of five stacked layers: execution, resource, model, agent, and application. Each lower layer supplies capabilities to the one above it: execution-layer optimizations make inference feasible on diverse hardware; resource-layer parallelism and scaling make the system elastic; model-layer compression and token reduction make large models lightweight enough for edge-cloud use; agent-layer components (multi-agent frameworks, planning, memory, tool use) turn the model into a service; application-layer batching and applications deliver the user-facing QoS. The paper's discovery is organizational rather than empirical: it claims these bodies of work belong to a single design space and that their integration, not any single technique, is the open research agenda.
Load-bearing premise
The load-bearing premise is that the techniques surveyed at different layers can meaningfully be integrated into one coherent deployment stack for agent services, and that the paper's selection of topics is representative enough to support its conclusions about open problems; the paper does not implement or demonstrate that the layers compose in practice.
Editorial extensions
If this is right
- If the framework is right, a serving system for FM agents should be designed with cross-layer budgets: a latency or accuracy target at the application layer should be traceable down to choices in execution, resource, and model layers.
- The survey's own lesson about agent-layer elasticity implies that future serving systems will need adaptive agents that decide when to call APIs, retrieve knowledge, or collaborate with other agents based on current load and task complexity.
- Edge-cloud deployment of large FMs becomes viable only when parallelism, model compression, and communication optimization are co-designed, since no single hardware class can host the full model.
- Multi-modal and mixture-of-experts models will require new serving-system mechanisms, because their activated modules and resource demands vary with the input.
- Batching and scheduling must become heterogeneity-aware, grouping requests by length, service characteristics, and per-request adapters rather than assuming a uniform model.
Reading between the lines
- Editorial inference: the layered framework implies a concrete design recipe — start from agent-level QoS requirements and derive lower-layer optimization targets — which the paper describes but does not itself validate end-to-end.
- Editorial inference: the identified agent-layer elasticity gap could be addressed by a scheduler that dynamically selects planning depth, tool-use rate, and collaboration topology under latency constraints; testing such a scheduler would be a natural next step beyond the survey.
- Editorial inference: cross-layer interactions may create non-compositional effects not quantified in the survey; for example, token reduction changes KV-cache size and attention patterns, which in turn shifts the optimal parallelism and batching strategy.
- Editorial inference: if the framework is accepted, a useful benchmark would be an open testbed that measures the same agent workload across different layer configurations, making the survey's taxonomy directly actionable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey proposes a five-layer unified framework for deploying foundation model (FM) powered agent services on heterogeneous edge-cloud devices. The layers are: execution optimization (hardware-specific computation, memory and communication optimizations), resource allocation and parallelism (cloud/edge resource scaling and various parallelism strategies), model-layer optimizations (model compression, quantization, distillation, token reduction), agent-layer components (multi-agent frameworks, planning, memory, tool use), and application-layer concerns (batching and representative applications). The paper reviews a large number of recent works in each area, provides several summary tables, and concludes with lessons learned and future research directions. Its central claim is that it is the first comprehensive survey of the deployment of real-time FM-powered agent services in heterogeneous devices.
Significance. The paper has clear value as a broad compilation of recent work on efficient FM inference and on AI-agent architectures. The taxonomy is readable and the tables (e.g., Table II on integrated frameworks, Table XIII on batching) provide useful entry points for readers. The paper is also honest in Section VII about current gaps, which is a strength. However, the significance of the paper as claimed depends entirely on whether the surveyed literature actually constitutes a coherent area of 'FM-powered agent service deployment.' The evidence in the manuscript does not support this: the systems in Sections II-IV are generic FM/LLM inference systems, and the agent work in Section V is algorithmic rather than deployment-oriented. The proposed framework is a plausible future integration agenda, but the claim of a first comprehensive survey of an existing research area is overstated. If the scope were reframed accordingly, the survey would still be a useful reference.
major comments (3)
- [I (Introduction, firstness claim) and Sections II-V] The paper's central claim that it is 'the first comprehensive survey to review and discuss the deployment of real-time FM-powered agent services in heterogeneous devices' (Introduction) is not supported by the material surveyed. None of the systems cited in Sections II-IV (e.g., FlashAttention, vLLM, PowerInfer, SpotServe, the resource-allocation works in Table III) is designed for or evaluated on agent workloads such as tool calling, planning loops, multi-agent coordination, or memory retrieval. Conversely, the agent frameworks in Section V (AgentVerse, Toolformer, DEPS, etc.) are described from an algorithmic/application perspective with no system-level deployment contributions. The paper itself concedes in Section VII-A3 that there is a 'significant gap in elasticity at the agent layer' and in Section VII-A1 that heterogeneous edge-cloud FM serving is 'under-explored.' Thus, on the evidence provided, the framework in Figure 1 is a proposal for future integration rather than a taxonomy of an existing body of work on agent-service deployment. This missing evidence is load-bearing because the firstness claim is the paper's principal contribution.
- [II-D and III] The connection between the surveyed infrastructure and agent services is never established. For example, Table II lists llama.cpp, MLC-LLM, FastChat, and similar frameworks, but the discussion does not explain how these frameworks support agent-specific requirements such as maintaining multi-turn tool-call state, sharing KV cache across planning iterations, or dynamically deciding when to offload subtasks to different devices. Likewise, the resource-allocation methods in Table III target generic DNN/LLM inference and do not model agent-specific request graphs, inter-agent communication, or memory-retrieval latency. The survey would need at least one worked example or a dedicated analysis showing how the layers compose for an agent service; without this, the unified framework is only a juxtaposition of two adjacent literatures.
- [VII-B3 and VII-C] Section VII-B3 lists 'Specific serving system for agents' as a future direction, acknowledging that current serving systems are designed for FM inference rather than for agent services. This is an honest statement, but it directly contradicts the introductory claim that the paper surveys the deployment of FM-powered agent services. The conclusion in Section VII-C repeats that the framework 'showcases the latest advancements' in this area, which is not supported by the content. The authors should either reframe the paper as a survey of building blocks plus a research agenda, or substantially expand the survey to include systems (if any exist) that actually address agent-service deployment end to end.
minor comments (3)
- [Throughout] There are many typographical errors and inconsistencies, including stray letters in the author affiliations (e.g., 'Y . Fan', 'V'), inconsistent use of backslashes in Table V, and a duplicated 'Section V' label in Figure 2 (the application layer should be Section VI).
- [I] The statistic on ChatGPT users is cited to a non-academic blog-style source (Nerdynav). A more authoritative source, such as a company report or a peer-reviewed citation, would be preferable for a survey.
- [V-C] The sentence beginning 'This rethinking method helps...' is grammatically incomplete, and the reference to Hatalis is cited without a first author name or paper title. Please ensure all citations are complete and the prose is polished.
Circularity Check
No circularity found: the survey organizes external literature; its firstness claim is a novelty assertion, not a derivation.
full rationale
This is a survey paper, not a derivation or empirical study. It contains no equations that are fitted to data, no parameters calibrated to a subset and then validated on a closely related quantity, and no predictive claims that reduce to its inputs by construction. The central claim is that the paper is the first comprehensive survey of deploying real-time FM-powered agent services in heterogeneous devices; this is a novelty and scope assertion about the literature, not a mathematical or statistical result derived from the framework. The unified framework in Figure 1 is an organizing taxonomy that maps existing work into execution, resource, model, agent, and application layers; such a taxonomy is a presentation device rather than a load-bearing inference. Sections II-IV review external systems for FM inference, compression, parallelism, and resource allocation, while Section V reviews external agent frameworks; none of these sections derives its conclusions from the survey's own framework. Self-citations, such as the prior edge-cloud AIGC survey [7], are used only to position the paper against previous surveys and do not carry the technical weight of the review; even if an author overlaps, the reviewed content consists of independently published systems and is externally checkable. The paper's own limitations, including the stated 'significant gap in elasticity at the agent layer' (Section VII-A3) and the admission that heterogeneous edge computing for FMs is 'under-explored' (Section VII-A1), honestly narrow the claimed scope, which may affect the strength of the firstness claim but is not a form of circularity. No passage defines the survey's key terms in terms of its conclusions, and no cited prior result is invoked as an unexamined theorem to force a choice. The result is therefore self-contained as a literature survey, and the correct circularity finding is no significant circularity.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Deploying Foundation Model Powered Agent Services: A Survey." pith.science (2026). https://pith.science/paper/62L5C7UK
@misc{pith2026241213437,
author = {Pith},
title = {Pith review of: Deploying Foundation Model Powered Agent Services: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/62L5C7UK}},
note = {Machine review of arXiv:2412.13437}
}
read the original abstract
Foundation model (FM) powered agent services are regarded as a promising solution to develop intelligent and personalized applications for advancing toward Artificial General Intelligence (AGI). To achieve high reliability and scalability in deploying these agent services, it is essential to collaboratively optimize computational and communication resources, thereby ensuring effective resource allocation and seamless service delivery. In pursuit of this vision, this paper proposes a unified framework aimed at providing a comprehensive survey on deploying FM-based agent services across heterogeneous devices, with the emphasis on the integration of model and resource optimization to establish a robust infrastructure for these services. Particularly, this paper begins with exploring various low-level optimization strategies during inference and studies approaches that enhance system scalability, such as parallelism techniques and resource scaling methods. The paper then discusses several prominent FMs and investigates research efforts focused on inference acceleration, including techniques such as model compression and token reduction. Moreover, the paper also investigates critical components for constructing agent services and highlights notable intelligent applications. Finally, the paper presents potential research directions for developing real-time agent services with high Quality of Service (QoS).
Figures
Figures from the paper (10 more)
Forward citations
Cited by 1 Pith paper
-
Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables
AL4RAG uses a retrieval-aware similarity metric to select annotation-worthy RAG conversation records, yielding DPO-trained models that reject hallucination-prone queries and preserve answer quality.
Reference graph
Works this paper leans on
-
[1]
On the Opportunities and Risks of Foundation Models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolf- sson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Dur- mus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K...
2022
-
[2]
The Rise and Potential of Large Language Model Based Agents: A Survey,
Z. Xi, W. Chen, X. Guo, W. He, Y . Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, R. Zheng, X. Fan, X. Wang, L. Xiong, Y . Zhou, W. Wang, C. Jiang, Y . Zou, X. Liu, Z. Yin, S. Dou, R. Weng, W. Cheng, Q. Zhang, W. Qin, Y . Zheng, X. Qiu, X. Huang, and T. Gui, “The Rise and Potential of Large Language Model Based Agents: A Survey,” 2023
2023
-
[3]
107 Up-to-Date ChatGPT Statistics & User Numbers,
Nerdynav, “107 Up-to-Date ChatGPT Statistics & User Numbers,” 2024, accessed: 2024-04-24. [Online]. Available: https://nerdynav. com/chatgpt-statistics/
2024
-
[4]
A Survey on Hardware Accelerators for Large Language Models,
C. Kachris, “A Survey on Hardware Accelerators for Large Language Models,” 2024
2024
-
[5]
A Survey on Scheduling Techniques in Computing and Network Convergence,
S. Tang, Y . Yu, H. Wang, G. Wang, W. Chen, Z. Xu, S. Guo, and W. Gao, “A Survey on Scheduling Techniques in Computing and Network Convergence,” IEEE Communications Surveys & Tutorials , vol. 26, no. 1, pp. 160–195, 2024
2024
-
[6]
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems,
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, H. Jin, T. Chen, and Z. Jia, “Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems,” 2023
2023
-
[7]
Unleashing the Power of Edge-Cloud Generative AI in Mobile Net- works: A Survey of AIGC Services,
M. Xu, H. Du, D. Niyato, J. Kang, Z. Xiong, S. Mao, Z. Han, A. Jamalipour, D. I. Kim, X. Shen, V . C. M. Leung, and H. V . Poor, “Unleashing the Power of Edge-Cloud Generative AI in Mobile Net- works: A Survey of AIGC Services,” IEEE Communications Surveys & Tutorials, pp. 1–1, 2024
2024
-
[8]
Machine and Deep Learning for Resource Allocation in Multi-Access Edge Computing: A Survey,
H. Djigal, J. Xu, L. Liu, and Y . Zhang, “Machine and Deep Learning for Resource Allocation in Multi-Access Edge Computing: A Survey,” IEEE Communications Surveys & Tutorials , vol. 24, no. 4, pp. 2449– 2494, 2022
2022
Show all 300 references
-
[9]
A Survey of Large Language Models,
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong, Y . Du, C. Yang, Y . Chen, Z. Chen, J. Jiang, R. Ren, Y . Li, X. Tang, Z. Liu, P. Liu, J.-Y . Nie, and J.-R. Wen, “A Survey of Large Language Models,” 2023
2023
-
[10]
A Comprehensive Overview of Large Language Models,
H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian, “A Comprehensive Overview of Large Language Models,” 2024
2024
-
[11]
A Survey on Model Compression for Large Language Models,
X. Zhu, J. Li, Y . Liu, C. Ma, and W. Wang, “A Survey on Model Compression for Large Language Models,” 2023
2023
-
[12]
Model Compression and Efficient Inference for Large Language Models: A Survey,
W. Wang, W. Chen, Y . Luo, Y . Long, Z. Lin, L. Zhang, B. Lin, D. Cai, and X. He, “Model Compression and Efficient Inference for Large Language Models: A Survey,” 2024
2024
-
[13]
A survey on knowledge distillation of large language models,
X. Xu, M. Li, C. Tao, T. Shen, R. Cheng, J. Li, C. Xu, D. Tao, and T. Zhou, “A survey on knowledge distillation of large language models,” arXiv preprint arXiv:2402.13116 , 2024
2024 arXiv
-
[14]
A survey on large language model based autonomous agents,
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y . Lin et al. , “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, pp. 1–26, 2024
2024
-
[15]
Large language model based multi-agents: A survey of progress and challenges,
T. Guo, X. Chen, Y . Wang, R. Chang, S. Pei, N. V . Chawla, O. Wiest, and X. Zhang, “Large language model based multi-agents: A survey of progress and challenges,” arXiv preprint arXiv:2402.01680 , 2024
2024 arXiv
-
[16]
FedDSE: Distribution-aware Sub-model Extraction for Fed- erated Learning over Resource-constrained Devices,
H. Wang, Y . Jia, M. Zhang, Q. Hu, H. Ren, P. Sun, Y . Wen, and T. Zhang, “FedDSE: Distribution-aware Sub-model Extraction for Fed- erated Learning over Resource-constrained Devices,” in Proceedings of the ACM on Web Conference 2024 , 2024, pp. 2902–2913
2024
-
[17]
Hardware accelerator for multi-head attention and position-wise feed-forward in the trans- former,
S. Lu, M. Wang, S. Liang, J. Lin, and Z. Wang, “Hardware accelerator for multi-head attention and position-wise feed-forward in the trans- former,” in 2020 IEEE 33rd International System-on-Chip Conference (SOCC). IEEE, 2020, pp. 84–89
2020
-
[18]
Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,
H. Jang, J. Kim, J.-E. Jo, J. Lee, and J. Kim, “Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,” in Proceedings of the 46th International Symposium on Computer Architecture, 2019, pp. 250–263
2019
-
[19]
Npe: An fpga-based overlay processor for natural language processing,
H. Khan, A. Khan, Z. Khan, L. B. Huang, K. Wang, and L. He, “Npe: An fpga-based overlay processor for natural language processing,” arXiv preprint arXiv:2104.06535 , 2021
2021 arXiv
-
[20]
Dfx: A low-latency multi-fpga appliance for accelerating transformer- based text generation,
S. Hong, S. Moon, J. Kim, S. Lee, M. Kim, D. Lee, and J.-Y . Kim, “Dfx: A low-latency multi-fpga appliance for accelerating transformer- based text generation,” in 2022 55th IEEE/ACM International Sympo- sium on Microarchitecture (MICRO) . IEEE, 2022, pp. 616–630
2022
-
[21]
Transformer- opu: An fpga-based overlay processor for transformer networks,
Y . Bai, H. Zhou, K. Zhao, J. Chen, J. Yu, and K. Wang, “Transformer- opu: An fpga-based overlay processor for transformer networks,” in 2023 IEEE 31st Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM) . IEEE, 2023, pp. 221–221
2023
-
[22]
A cost-efficient fpga imple- mentation of tiny transformer model using neural ode,
I. Okubo, K. Sugiura, and H. Matsutani, “A cost-efficient fpga imple- mentation of tiny transformer model using neural ode,” arXiv preprint arXiv:2401.02721, 2024
2024 arXiv
-
[23]
Flightllm: Efficient large language model inference with a complete mapping flow on fpga,
S. Zeng, J. Liu, G. Dai, X. Yang, T. Fu, H. Wang, W. Ma, H. Sun, S. Li, Z. Huang et al. , “Flightllm: Efficient large language model inference with a complete mapping flow on fpga,” arXiv preprint arXiv:2401.03868, 2024
2024 arXiv
-
[24]
Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,
T. J. Ham, S. J. Jung, S. Kim, Y . H. Oh, Y . Park, Y . Song, J.-H. Park, S. Lee, K. Park, J. W. Lee et al. , “Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA)....
2020
-
[25]
Elsa: Hardware-software co-design for efficient, lightweight self- attention mechanism in neural networks,
T. J. Ham, Y . Lee, S. H. Seo, S. Kim, H. Choi, S. J. Jung, and J. W. Lee, “Elsa: Hardware-software co-design for efficient, lightweight self- attention mechanism in neural networks,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) . IEEE, ...
2021
-
[26]
Spatten: Efficient sparse attention architecture with cascade token and head pruning,
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2021, pp. 97–110
2021
-
[27]
Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,
L. Lu, Y . Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y . Liang, “Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture, 2021, pp. 977–991
2021
-
[28]
Energon: Toward efficient acceleration of transformers using dynamic sparse attention,
Z. Zhou, J. Liu, Z. Gu, and G. Sun, “Energon: Toward efficient acceleration of transformers using dynamic sparse attention,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 42, no. 1, pp. 136–149, 2022
2022
-
[29]
Att: A fault-tolerant reram accelerator for attention-based neural networks,
H. Guo, L. Peng, J. Zhang, Q. Chen, and T. D. LeCompte, “Att: A fault-tolerant reram accelerator for attention-based neural networks,” in 2020 IEEE 38th International Conference on Computer Design (ICCD). IEEE, 2020, pp. 213–221
2020
-
[30]
In-memory com- puting based accelerator for transformer networks for long sequences,
A. F. Laguna, A. Kazemi, M. Niemier, and X. S. Hu, “In-memory com- puting based accelerator for transformer networks for long sequences,” in 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 2021, pp. 1839–1844
2021
-
[31]
Work in progress: Real-time transformer inference on edge ai accelerators,
B. Reidy, M. Mohammadi, M. Elbtity, H. Smith, and Z. Ramtin, “Work in progress: Real-time transformer inference on edge ai accelerators,” in 2023 IEEE 29th Real-Time and Embedded Technology and Appli- cations Symposium (RTAS) , 2023, pp. 341–344
2023
-
[32]
Simplifying transformer blocks,
B. He and T. Hofmann, “Simplifying transformer blocks,” arXiv preprint arXiv:2311.01906, 2023
2023 arXiv
-
[33]
Accelerating transformer networks through recomposing softmax layers,
J. Choi, H. Li, B. Kim, S. Hwang, and J. H. Ahn, “Accelerating transformer networks through recomposing softmax layers,” in 2022 IEEE International Symposium on Workload Characterization (IISWC). IEEE, 2022, pp. 92–103
2022
-
[34]
Inference with reference: Lossless acceleration of large language models,
N. Yang, T. Ge, L. Wang, B. Jiao, D. Jiang, L. Yang, R. Majumder, and F. Wei, “Inference with reference: Lossless acceleration of large language models,” arXiv preprint arXiv:2304.04487 , 2023
2023 arXiv
-
[35]
Exponentially faster language mod- elling,
P. Belcak and R. Wattenhofer, “Exponentially faster language mod- elling,” arXiv preprint arXiv:2311.10770 , 2023
2023 arXiv
-
[36]
Efficient llm inference on cpus,
H. Shen, H. Chang, B. Dong, Y . Luo, and H. Meng, “Efficient llm inference on cpus,” arXiv preprint arXiv:2311.00502 , 2023
2023 arXiv
-
[37]
Powerinfer: Fast large language model serving with a consumer-grade gpu,
Y . Song, Z. Mi, H. Xie, and H. Chen, “Powerinfer: Fast large language model serving with a consumer-grade gpu,” arXiv preprint arXiv:2312.12456, 2023
2023 arXiv
-
[38]
Hetegen: Het- erogeneous parallel inference for large language models on resource- constrained devices,
X. Zhao, B. Jia, H. Zhou, Z. Liu, S. Cheng, and Y . You, “Hetegen: Het- erogeneous parallel inference for large language models on resource- constrained devices,” arXiv preprint arXiv:2403.01164 , 2024
2024 arXiv
-
[39]
Flexgen: High-throughput generative inference of large language models with a single gpu,
Y . Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Re, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single gpu,” 2023
2023
-
[40]
Deja vu: Contextual sparsity MANUSCRIPT 35 for efficient llms at inference time,
Z. Liu, J. Wang, T. Dao, T. Zhou, B. Yuan, Z. Song, A. Shrivastava, C. Zhang, Y . Tian, C. Re et al. , “Deja vu: Contextual sparsity MANUSCRIPT 35 for efficient llms at inference time,” in International Conference on Machine Learning. PMLR, 2023, pp. 22 137–22 176
2023
-
[41]
Llm in a flash: Efficient large language model inference with limited memory,
K. Alizadeh, I. Mirzadeh, D. Belenko, K. Khatamifard, M. Cho, C. C. Del Mundo, M. Rastegari, and M. Farajtabar, “Llm in a flash: Efficient large language model inference with limited memory,” arXiv preprint arXiv:2312.11514, 2023
2023 arXiv
-
[42]
Fast transformer decoding: One write-head is all you need,
N. Shazeer, “Fast transformer decoding: One write-head is all you need,” arXiv preprint arXiv:1911.02150 , 2019
1911 arXiv
-
[43]
Gqa: Training generalized multi-query transformer models from multi-head checkpoints,
J. Ainslie, J. Lee-Thorp, M. de Jong, Y . Zemlyanskiy, F. Lebron, and S. Sanghai, “Gqa: Training generalized multi-query transformer models from multi-head checkpoints,” in The 2023 Conference on Empirical Methods in Natural Language Processing , 2023
2023
-
[44]
Efficient memory management for large language model serving with pagedattention,
W. Kwon, Z. Li, S. Zhuang, Y . Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
-
[45]
Flashattention: Fast and memory-efficient exact attention with io-awareness,
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. R ´e, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022
2022
-
[46]
Flashattention-2: Faster attention with better parallelism and work partitioning,
T. Dao, “Flashattention-2: Faster attention with better parallelism and work partitioning,” arXiv preprint arXiv:2307.08691 , 2023
2023 arXiv
-
[47]
Flashdecoding++: Faster large language model inference on gpus,
K. Hong, G. Dai, J. Xu, Q. Mao, X. Li, J. Liu, K. Chen, H. Dong, and Y . Wang, “Flashdecoding++: Faster large language model inference on gpus,” arXiv preprint arXiv:2311.01282 , 2023
2023 arXiv
-
[48]
Bminf: An efficient toolkit for big model inference and tuning,
X. Han, G. Zeng, W. Zhao, Z. Liu, Z. Zhang, J. Zhou, J. Zhang, J. Chao, and M. Sun, “Bminf: An efficient toolkit for big model inference and tuning,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, 2022, pp. 224– 230
2022
-
[49]
Splitwise: Efficient generative llm inference using phase splitting,
P. Patel, E. Choukse, C. Zhang, ´I. Goiri, A. Shah, S. Maleki, and R. Bianchini, “Splitwise: Efficient generative llm inference using phase splitting,” arXiv preprint arXiv:2311.18677 , 2023
2023 arXiv
-
[50]
Fast distributed inference serving for large language models,
B. Wu, Y . Zhong, Z. Zhang, G. Huang, X. Liu, and X. Jin, “Fast distributed inference serving for large language models,” arXiv preprint arXiv:2305.05920, 2023
2023 arXiv
-
[51]
Specinfer: Accelerating generative large language model serving with speculative inference and token tree verification,
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, R. Y . Y . Wong, A. Zhu, L. Yang, X. Shi, C. Shi, Z. Chen, D. Arfeen, R. Abhyankar, and Z. Jia, “Specinfer: Accelerating generative large language model serving with speculative inference and token tree verification,” 2023
2023
-
[52]
Llmcad: Fast and scalable on-device large language model inference,
D. Xu, W. Yin, X. Jin, Y . Zhang, S. Wei, M. Xu, and X. Liu, “Llmcad: Fast and scalable on-device large language model inference,” arXiv preprint arXiv:2309.04255, 2023
2023 arXiv
-
[53]
Deepspeed-moe: Advancing mixture- of-experts inference and training to power next-generation ai scale,
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y . Aminabadi, A. A. Awan, J. Rasley, and Y . He, “Deepspeed-moe: Advancing mixture- of-experts inference and training to power next-generation ai scale,” in International conference on machine learning . PMLR, 2022, pp. 18 332–18 346
2022
-
[54]
Edgemoe: Fast on-device inference of moe-based large language models,
R. Yi, L. Guo, S. Wei, A. Zhou, S. Wang, and M. Xu, “Edgemoe: Fast on-device inference of moe-based large language models,” arXiv preprint arXiv:2308.14352, 2023
2023 arXiv
-
[55]
Fast inference of mixture-of-experts lan- guage models with offloading,
A. Eliseev and D. Mazur, “Fast inference of mixture-of-experts lan- guage models with offloading,”arXiv preprint arXiv:2312.17238, 2023
2023 arXiv
-
[56]
Moe-infinity: Activation- aware expert offloading for efficient moe serving,
L. Xue, Y . Fu, Z. Lu, L. Mai, and M. Marina, “Moe-infinity: Activation- aware expert offloading for efficient moe serving,” arXiv preprint arXiv:2401.14361, 2024
2024 arXiv
-
[57]
Fiddler: Cpu-gpu orchestration for fast inference of mixture-of-experts models,
K. Kamahori, Y . Gu, K. Zhu, and B. Kasikci, “Fiddler: Cpu-gpu orchestration for fast inference of mixture-of-experts models,” arXiv preprint arXiv:2402.07033, 2024
2024 arXiv
-
[58]
Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference,
R. Hwang, J. Wei, S. Cao, C. Hwang, X. Tang, T. Cao, M. Yang, and M. Rhu, “Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference,” arXiv preprint arXiv:2308.12066, 2023
2023 arXiv
-
[59]
Intelligence-endogenous management platform for computing and network convergence,
Z. Hong, X. Qiu, J. Lin, W. Chen, Y . Yu, H. Wang, S. Guo, and W. Gao, “Intelligence-endogenous management platform for computing and network convergence,” IEEE Network, 2023
2023
-
[60]
Resource allocation in large language model integrated 6g vehicular networks,
C. Liu and J. Zhao, “Resource allocation in large language model integrated 6g vehicular networks,” arXiv preprint arXiv:2403.19016 , 2024
2024 arXiv
-
[61]
Lingualinked: A distributed large language model inference system for mobile de- vices,
J. Zhao, Y . Song, S. Liu, I. G. Harris, and S. A. Jyothi, “Lingualinked: A distributed large language model inference system for mobile de- vices,” arXiv preprint arXiv:2312.00388 , 2023
2023 arXiv
-
[62]
{MegaScale}: Scaling large language model training to more than 10,000 {GPUs},
Z. Jiang, H. Lin, Y . Zhong, Q. Huang, Y . Chen, Z. Zhang, Y . Peng, X. Li, C. Xie, S. Nong et al. , “ {MegaScale}: Scaling large language model training to more than 10,000 {GPUs},” in 21st USENIX Sym- posium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp...
2024
-
[63]
LOSP: Overlap synchronization parallel with local compensation for fast dis- tributed training,
H. Wang, Z. Qu, S. Guo, N. Wang, R. Li, and W. Zhuang, “LOSP: Overlap synchronization parallel with local compensation for fast dis- tributed training,” IEEE Journal on Selected Areas in Communications, vol. 39, no. 8, pp. 2541–2557, 2021
2021
-
[64]
Heterogeneous semantic and bit communications: A semi-noma scheme,
X. Mu, Y . Liu, L. Guo, and N. Al-Dhahir, “Heterogeneous semantic and bit communications: A semi-noma scheme,” IEEE Journal on Selected Areas in Communications , vol. 41, no. 1, pp. 155–169, 2022
2022
-
[65]
Computing networks enabled semantic communications,
Z. Qin, J. Ying, D. Yang, H. Wang, and X. Tao, “Computing networks enabled semantic communications,” IEEE Network, 2024
2024
-
[66]
ggerganov/llama.cpp: Port of facebook’s llama model in c/c++
G. Gerganov, “ggerganov/llama.cpp: Port of facebook’s llama model in c/c++.” https://github.com/ggerganov/llama.cpp, 2023
2023
-
[67]
MLC-LLM,
M. team, “MLC-LLM,” 2023. [Online]. Available: https://github.com/ mlc-ai/mlc-llm
2023
-
[68]
mnn-llm: llm deploy project based mnn
mnn llm, “mnn-llm: llm deploy project based mnn.” https://github.com/ wangzhaode/mnn-llm, 2023
2023
-
[69]
Judging llm-as-a-judge with mt-bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023
2023
-
[70]
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,
R. Y . Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al., “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Com- ...
2022
-
[71]
Openvino deep learning workbench: Comprehensive analysis and tuning of neural networks inference,
Y . Gorbachev, M. Fedorov, I. Slavutin, A. Tugarev, M. Fatekhov, and Y . Tarkan, “Openvino deep learning workbench: Comprehensive analysis and tuning of neural networks inference,” inProceedings of the IEEE/CVF International Conference on Computer Vision Workshops , 2019, pp. 0–0
2019
-
[72]
mllm is a fast and lightweight multimodal llm inference engine for mobile and edge devices
mllm, “mllm is a fast and lightweight multimodal llm inference engine for mobile and edge devices.” https://github.com/UbiquitousLearning/ mllm, 2023
2023
-
[73]
Fp6-llm: Efficiently serving large language models through fp6-centric algorithm-system co-design,
H. Xia, Z. Zheng, X. Wu, S. Chen, Z. Yao, S. Youn, A. Bakhtiari, M. Wyatt, D. Zhuang, Z. Zhouet al., “Fp6-llm: Efficiently serving large language models through fp6-centric algorithm-system co-design,” arXiv preprint arXiv:2401.14112 , 2024
2024 arXiv
-
[74]
Colossal-ai: A unified deep learning system for large-scale parallel training,
S. Li, H. Liu, Z. Bian, J. Fang, H. Huang, Y . Liu, B. Wang, and Y . You, “Colossal-ai: A unified deep learning system for large-scale parallel training,” in Proceedings of the 52nd International Conference on Parallel Processing, 2023, pp. 766–775
2023
-
[75]
Efficient large-scale language model training on gpu clusters using megatron-lm,
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V . Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro et al. , “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference fo...
2021
-
[76]
A tensorrt toolbox for optimized large language model inference,
tensorrtllm, “A tensorrt toolbox for optimized large language model inference,” https://github.com/NVIDIA/TensorRT-LLM, 2023
2023
-
[77]
Langchain,
Harrison Chase, “Langchain,” https://github.com/langchain-ai/ langchain, 2024, accessed: 2024-04-07
2024
-
[78]
Parrot: Efficient Serving of LLM-based Applications with Semantic Variable,
C. Lin, Z. Han, C. Zhang, Y . Yang, F. Yang, C. Chen, and L. Qiu, “Parrot: Efficient Serving of LLM-based Applications with Semantic Variable,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara, CA: USENIX Associa- tion, Jul. 2024
2024
-
[79]
SGLang: Efficient Execution of Structured Language Model Pro- grams,
L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y . Sheng, “SGLang: Efficient Execution of Structured Language Model Pro- grams,” 2024
2024
-
[80]
Clipper: A {Low-Latency} online prediction serving system,
D. Crankshaw, X. Wang, G. Zhou, M. J. Franklin, J. E. Gonzalez, and I. Stoica, “Clipper: A {Low-Latency} online prediction serving system,” in 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) , 2017, pp. 613–627
2017
-
[81]
{MArk}: Exploiting cloud services for {Cost-Effective},{SLO-Aware} machine learning infer- ence serving,
C. Zhang, M. Yu, W. Wang, and F. Yan, “ {MArk}: Exploiting cloud services for {Cost-Effective},{SLO-Aware} machine learning infer- ence serving,” in 2019 USENIX Annual Technical Conference (USENIX ATC 19), 2019, pp. 1049–1062
2019
-
[82]
Nexus: A GPU cluster engine for accelerating DNN-based video analysis,
H. Shen, L. Chen, Y . Jin, L. Zhao, B. Kong, M. Philipose, A. Kr- ishnamurthy, and R. Sundaram, “Nexus: A GPU cluster engine for accelerating DNN-based video analysis,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles, 2019, pp. 322–337
2019
-
[83]
InferLine: latency-aware provisioning and scaling for prediction serving pipelines,
D. Crankshaw, G.-E. Sela, X. Mo, C. Zumar, I. Stoica, J. Gonzalez, and A. Tumanov, “InferLine: latency-aware provisioning and scaling for prediction serving pipelines,” in Proceedings of the 11th ACM Symposium on Cloud Computing , 2020, pp. 477–491
2020
-
[84]
Serving {DNNs} like clockwork: Performance predictability from the bottom up,
A. Gujarati, R. Karimi, S. Alzayat, W. Hao, A. Kaufmann, Y . Vig- fusson, and J. Mace, “Serving {DNNs} like clockwork: Performance predictability from the bottom up,” in 14th USENIX Symposium on MANUSCRIPT 36 Operating Systems Design and Implementation (OSDI 20) , 2020, pp. 443–462
2020
-
[85]
{INFaaS}: Automated model-less inference serving,
F. Romero, Q. Li, N. J. Yadwadkar, and C. Kozyrakis, “ {INFaaS}: Automated model-less inference serving,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21) , 2021, pp. 397–411
2021
-
[86]
Morphling: Fast, near-optimal auto-configuration for cloud- native model serving,
L. Wang, L. Yang, Y . Yu, W. Wang, B. Li, X. Sun, J. He, and L. Zhang, “Morphling: Fast, near-optimal auto-configuration for cloud- native model serving,” inProceedings of the ACM Symposium on Cloud Computing, 2021, pp. 639–653
2021
-
[87]
Cocktail: A multidimensional optimization for model serving in cloud,
Gunasekaran, Jashwant Raj and Mishra, Cyan Subhra and Thinakaran, Prashanth and Sharma, Bikash and Kandemir, Mahmut Taylan and Das, Chita R, “Cocktail: A multidimensional optimization for model serving in cloud,” in 19th USENIX Symposium on Networked Systems Design and Impleme...
2022
-
[88]
Kairos: Building cost- efficient machine learning inference systems with heterogeneous cloud resources,
B. Li, S. Samsi, V . Gadepally, and D. Tiwari, “Kairos: Building cost- efficient machine learning inference systems with heterogeneous cloud resources,” in Proceedings of the 32nd International Symposium on High-Performance Parallel and Distributed Computing , 2023, pp. 3– 16
2023
-
[89]
{SHEPHERD}: Serving {DNNs} in the wild,
H. Zhang, Y . Tang, A. Khandelwal, and I. Stoica, “ {SHEPHERD}: Serving {DNNs} in the wild,” in 20th USENIX Symposium on Net- worked Systems Design and Implementation (NSDI 23), 2023, pp. 787– 808
2023
-
[90]
Spotserve: Serving generative large language models on preemptible instances,
X. Miao, C. Shi, J. Duan, X. Xi, D. Lin, B. Cui, and Z. Jia, “Spotserve: Serving generative large language models on preemptible instances,” arXiv preprint arXiv:2311.15566 , 2023
2023 arXiv
-
[91]
Frequency resource allocation and interference management in mobile edge com- puting for an Internet of Things system,
W. Na, S. Jang, Y . Lee, L. Park, N.-N. Dao, and S. Cho, “Frequency resource allocation and interference management in mobile edge com- puting for an Internet of Things system,” IEEE Internet of Things Journal, vol. 6, no. 3, pp. 4910–4920, 2018
2018
-
[92]
Decentralized resource auctioning for latency-sensitive edge computing,
C. Avasalcai, C. Tsigkanos, and S. Dustdar, “Decentralized resource auctioning for latency-sensitive edge computing,” in 2019 IEEE inter- national conference on edge computing (EDGE) . IEEE, 2019, pp. 72–76
2019
-
[93]
Joint computation partitioning and resource allocation for latency sensitive applications in mobile edge clouds,
L. Yang, B. Liu, J. Cao, Y . Sahni, and Z. Wang, “Joint computation partitioning and resource allocation for latency sensitive applications in mobile edge clouds,” IEEE Transactions on Services Computing , vol. 14, no. 5, pp. 1439–1452, 2019
2019
-
[94]
Adaptive computation offloading and resource allocation strategy in a mobile edge computing environment,
Z. Tong, X. Deng, F. Ye, S. Basodi, X. Xiao, and Y . Pan, “Adaptive computation offloading and resource allocation strategy in a mobile edge computing environment,”Information Sciences, vol. 537, pp. 116– 131, 2020
2020
-
[95]
Resource allocation based on deep reinforcement learning in IoT edge computing,
X. Xiong, K. Zheng, L. Lei, and L. Hou, “Resource allocation based on deep reinforcement learning in IoT edge computing,” IEEE Journal on Selected Areas in Communications , vol. 38, no. 6, pp. 1133–1146, 2020
2020
-
[96]
CE-IoT: Cost-effective cloud- edge resource provisioning for heterogeneous IoT applications,
Z. Zhou, S. Yu, W. Chen, and X. Chen, “CE-IoT: Cost-effective cloud- edge resource provisioning for heterogeneous IoT applications,” IEEE Internet of Things Journal , vol. 7, no. 9, pp. 8600–8614, 2020
2020
-
[97]
Dynamic resource allocation and computation offloading for IoT fog computing system,
Z. Chang, L. Liu, X. Guo, and Q. Sheng, “Dynamic resource allocation and computation offloading for IoT fog computing system,” IEEE Transactions on Industrial Informatics , vol. 17, no. 5, pp. 3348–3357, 2020
2020
-
[98]
Lass: Running latency sensitive serverless computations at the edge,
B. Wang, A. Ali-Eldin, and P. Shenoy, “Lass: Running latency sensitive serverless computations at the edge,” in Proceedings of the 30th international symposium on high-performance parallel and distributed computing, 2021, pp. 239–251
2021
-
[99]
Resource provisioning and allocation in function- as-a-service edge-clouds,
O. Ascigil, A. G. Tasiopoulos, T. K. Phan, V . Sourlas, I. Psaras, and G. Pavlou, “Resource provisioning and allocation in function- as-a-service edge-clouds,” IEEE Transactions on Services Computing , vol. 15, no. 4, pp. 2410–2424, 2021
2021
-
[100]
KneeScale: Efficient resource scaling for serverless computing at the edge,
X. Li, P. Kang, J. Molone, W. Wang, and P. Lama, “KneeScale: Efficient resource scaling for serverless computing at the edge,” in2022 22nd IEEE International Symposium on Cluster, Cloud and Internet Computing (CCGrid). IEEE, 2022, pp. 180–189
2022
-
[101]
CEC: A containerized edge computing framework for dynamic resource provisioning,
S. Hu, W. Shi, and G. Li, “CEC: A containerized edge computing framework for dynamic resource provisioning,” IEEE Transactions on Mobile Computing, 2022
2022
-
[102]
FedCDA: Federated Learning with Cross-rounds Divergence-aware Aggregation,
H. Wang, H. Xu, Y . Li, Y . Xu, R. Li, and T. Zhang, “FedCDA: Federated Learning with Cross-rounds Divergence-aware Aggregation,” in The Twelfth International Conference on Learning Representations , 2024
2024
-
[103]
Megatron-lm: Training multi-billion parameter language models using model parallelism,
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catan- zaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” arXiv preprint arXiv:1909.08053 , 2019
1909 arXiv
-
[104]
Efficiently scaling transformer inference,
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
-
[105]
Alpaserve: Statistical multiplexing with model parallelism for deep learning serving,
Z. Li, L. Zheng, Y . Zhong, V . Liu, Y . Sheng, X. Jin, Y . Huang, Z. Chen, H. Zhang, J. E. Gonzalez et al. , “Alpaserve: Statistical multiplexing with model parallelism for deep learning serving,” in 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2...
2023
-
[106]
Lightseq: Sequence level parallelism for distributed training of long context transformers,
D. Li, R. Shao, A. Xie, E. P. Xing, J. E. Gonzalez, I. Stoica, X. Ma, and H. Zhang, “Lightseq: Sequence level parallelism for distributed training of long context transformers,” arXiv preprint arXiv:2310.03294, 2023
2023 arXiv
-
[107]
Distributed Inference and Fine-tuning of Large Language Models Over The In- ternet,
A. Borzunov, M. Ryabinin, A. Chumachenko, D. Baranchuk, T. Dettmers, Y . Belkada, P. Samygin, and C. A. Raffel, “Distributed Inference and Fine-tuning of Large Language Models Over The In- ternet,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[108]
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills,
A. Agrawal, A. Panwar, J. Mohan, N. Kwatra, B. S. Gulavani, and R. Ramjee, “Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills,” arXiv preprint arXiv:2308.16369 , 2023
2023 arXiv
-
[109]
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving,
Y . Zhong, S. Liu, J. Chen, J. Hu, Y . Zhu, X. Liu, X. Jin, and H. Zhang, “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving,” arXiv preprint arXiv:2401.09670, 2024
2024 arXiv
-
[110]
Distributing deep neural networks with containerized partitions at the edge,
L. Zhou, H. Wen, R. Teodorescu, and D. H. Du, “Distributing deep neural networks with containerized partitions at the edge,” in 2nd USENIX Workshop on Hot Topics in Edge Computing (HotEdge 19) , 2019
2019
-
[111]
Distributed inference acceleration with adaptive DNN partitioning and offloading,
T. Mohammed, C. Joe-Wong, R. Babbar, and M. Di Francesco, “Distributed inference acceleration with adaptive DNN partitioning and offloading,” in IEEE INFOCOM 2020-IEEE Conference on Computer Communications. IEEE, 2020, pp. 854–863
2020
-
[112]
Coedge: Cooperative dnn inference with adaptive workload partitioning over heterogeneous edge devices,
L. Zeng, X. Chen, Z. Zhou, L. Yang, and J. Zhang, “Coedge: Cooperative dnn inference with adaptive workload partitioning over heterogeneous edge devices,” IEEE/ACM Transactions on Networking, vol. 29, no. 2, pp. 595–608, 2020
2020
-
[113]
Throughput maximization of delay-aware DNN inference in edge computing by exploring DNN model partitioning and inference parallelism,
J. Li, W. Liang, Y . Li, Z. Xu, X. Jia, and S. Guo, “Throughput maximization of delay-aware DNN inference in edge computing by exploring DNN model partitioning and inference parallelism,” IEEE Transactions on Mobile Computing , vol. 22, no. 5, pp. 3017–3030, 2021
2021
-
[114]
Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures,
Y . Wu, M. Lentz, D. Zhuo, and Y . Lu, “Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures,” Pro- ceedings of the VLDB Endowment , vol. 16, no. 3, pp. 406–419, 2022
2022
-
[115]
PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices,
Y . Hu, C. Imes, X. Zhao, S. Kundu, P. A. Beerel, S. P. Crago, and J. P. Walters, “PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices,” in 2022 25th Euromicro Conference on Digital System Design (DSD) , 2022, pp. 298–307
2022
-
[116]
PDD: partitioning DAG-topology DNNs for streaming tasks,
L. Wu, G. Gao, J. Yu, F. Zhou, Y . Yang, and T. Wang, “PDD: partitioning DAG-topology DNNs for streaming tasks,” IEEE Internet of Things Journal , 2023
2023
-
[117]
Dnn partitioning for inference throughput acceleration at the edge,
T. Feltin, L. March ´o, J.-A. Cordero-Fuertes, F. Brockners, and T. H. Clausen, “Dnn partitioning for inference throughput acceleration at the edge,” IEEE Access, vol. 11, pp. 52 236–52 249, 2023
2023
-
[118]
Distributed DNN Inference with Fine-grained Model Partitioning in Mobile Edge Computing Networks,
H. Li, X. Li, Q. Fan, Q. He, X. Wang, and V . C. Leung, “Distributed DNN Inference with Fine-grained Model Partitioning in Mobile Edge Computing Networks,” IEEE Transactions on Mobile Computing , 2024
2024
-
[119]
MoEI: Mobility-Aware Edge Inference Based on Model Partition and Service Migration,
Z. Liu, M. Tian, M. Dong, X. Wang, C. Qiu, and C. Zhang, “MoEI: Mobility-Aware Edge Inference Based on Model Partition and Service Migration,” IEEE Transactions on Mobile Computing, no. 01, pp. 1–14, 2024
2024
-
[120]
NVIDIA Triton Inference Server,
NVIDIA Corporation, “NVIDIA Triton Inference Server,” https: //developer.nvidia.com/nvidia-triton-inference-server, 2024, accessed: 2024-04-17
2024
-
[121]
TensorFlow Serving,
Google LLC, “TensorFlow Serving,” https://www.tensorflow.org/tfx/ guide/serving, 2024, accessed: 2024-04-17
2024
-
[122]
FedNLR: Federated Learning with Neuron-wise Learning Rates,
H. Wang, P. Zheng, X. Han, W. Xu, R. Li, and T. Zhang, “FedNLR: Federated Learning with Neuron-wise Learning Rates,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024, pp. 3069–3080
2024
-
[123]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[124]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research, vol. 21, no. 140, pp. 1–67, 2020. MANUSCRIPT 37
2020
-
[125]
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners.”
-
[126]
Language mod- els are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language mod- els are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020
1901
-
[127]
Pangu- α: Large-scale autoregressive pretrained chinese language models with auto-parallel computation,
W. Zeng, X. Ren, T. Su, H. Wang, Y . Liao, Z. Wang, X. Jiang, Z. Yang, K. Wang, X. Zhang et al. , “Pangu- α: Large-scale autoregressive pretrained chinese language models with auto-parallel computation,” arXiv e-prints, pp. arXiv–2104, 2021
2021
-
[128]
Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation,
Y . Sun, S. Wang, S. Feng, S. Ding, C. Pang, J. Shang, J. Liu, X. Chen, Y . Zhao, Y . Lu et al. , “Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation,” arXiv preprint arXiv:2107.02137, 2021
2021 arXiv
-
[129]
Jurassic-1: Technical details and evaluation,
O. Lieber, O. Sharir, B. Lenz, and Y . Shoham, “Jurassic-1: Technical details and evaluation,” White Paper. AI21 Labs, vol. 1, p. 9, 2021
2021
-
[130]
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clarket al., “Training compute-optimal large language models,” arXiv preprint arXiv:2203.15556, 2022
2022 arXiv
-
[131]
Lamda: Language models for dialog applications,
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y . Duet al., “Lamda: Language models for dialog applications,” arXiv preprint arXiv:2201.08239 , 2022
2022 arXiv
-
[132]
Us- ing deepspeed and megatron to train megatron-turing nlg 530b, a large- scale generative language model,
S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V . Korthikantiet al., “Us- ing deepspeed and megatron to train megatron-turing nlg 530b, a large- scale generative language model,” arXiv preprint arXiv:2201.11990 , 2022
2022 arXiv
-
[133]
Scaling language models: Methods, analysis & insights from training gopher,
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young et al., “Scaling language models: Methods, analysis & insights from training gopher,” arXiv preprint arXiv:2112.11446, 2021
2021 arXiv
-
[134]
Opt: Open pre-trained transformer language models,
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V . Lin et al. , “Opt: Open pre-trained transformer language models,” arXiv preprint arXiv:2205.01068, 2022
2022 arXiv
-
[135]
Galactica: A large language model for science,
R. Taylor, M. Kardas, G. Cucurull, T. Scialom, A. Hartshorn, E. Sar- avia, A. Poulton, V . Kerkez, and R. Stojnic, “Galactica: A large language model for science,” arXiv preprint arXiv:2211.09085 , 2022
2022 arXiv
-
[136]
Bloom: A 176b-parameter open-access multilingual language model,
T. Le Scao, A. Fan, C. Akiki, E. Pavlick, S. Ili ´c, D. Hesslow, R. Castagn ´e, A. S. Luccioni, F. Yvon, M. Gall ´e et al. , “Bloom: A 176b-parameter open-access multilingual language model,” 2023
2023
-
[137]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[138]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al. , “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[139]
Baichuan-7b: About: A large-scale 7b pretrain- ing language model developed by baichuan-inc,
BaiChuan-Inc, “Baichuan-7b: About: A large-scale 7b pretrain- ing language model developed by baichuan-inc,” https://github.com/ baichuan-inc/Baichuan-7B/tree/main, 2024, accessed: 2024-04-07
2024
-
[140]
A 13b large language model developed by baichuan intelligent technology,
Baichuan Intelligent Technology, “A 13b large language model developed by baichuan intelligent technology,” https://github.com/ baichuan-inc/Baichuan-13B/tree/main, 2024, accessed: 2024-04-07
2024
-
[141]
Qwen technical report,
J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huang et al. , “Qwen technical report,” arXiv preprint arXiv:2309.16609, 2023
2023 arXiv
-
[142]
Skywork: A more open bilingual foundation model,
T. Wei, L. Zhao, L. Zhang, B. Zhu, L. Wang, H. Yang, B. Li, C. Cheng, W. L ¨u, R. Hu et al. , “Skywork: A more open bilingual foundation model,” arXiv preprint arXiv:2310.19341 , 2023
2023 arXiv
-
[143]
The falcon series of open language models,
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojo- caru, M. Debbah, ´E. Goffinet, D. Hesslow, J. Launay, Q. Malartic et al. , “The falcon series of open language models,” arXiv preprint arXiv:2311.16867, 2023
2023 arXiv
-
[144]
Starcoder: may the source be with you!
R. Li, L. B. Allal, Y . Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim et al., “Starcoder: may the source be with you!” arXiv preprint arXiv:2305.06161 , 2023
2023 arXiv
-
[145]
Yi: Open foundation models by 01. ai,
A. Young, B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, H. Li, J. Zhu, J. Chen, J. Chang et al., “Yi: Open foundation models by 01. ai,” arXiv preprint arXiv:2403.04652 , 2024
2024 arXiv
-
[146]
Finetuned language models are zero-shot learners,
J. Wei, M. Bosma, V . Y . Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V . Le, “Finetuned language models are zero-shot learners,” 2022
2022
-
[147]
mt5: A massively multilingual pre-trained text-to-text transformer,
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,” arXiv preprint arXiv:2010.11934 , 2020
2010 arXiv
-
[148]
Flan-moe: Scaling instruction- finetuned language models with sparse mixture of experts,
S. Shen, L. Hou, Y . Zhou, N. Du, S. Longpre, J. Wei, H. W. Chung, B. Zoph, W. Fedus, X. Chen et al. , “Flan-moe: Scaling instruction- finetuned language models with sparse mixture of experts,” arXiv e- prints, pp. arXiv–2305, 2023
2023
-
[149]
Scaling instruction- finetuned language models,
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y . Tay, W. Fedus, Y . Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V . Zhao, Y . Huang, A. Dai, H. Y...
2022
-
[150]
Taori, I
R. Taori, I. Gulrajani, T. Zhang, Y . Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Alpaca,” 2023. [Online]. Available: https://crfm.stanford.edu/2023/03/13/alpaca.html
2023
-
[151]
Glm-130b: An open bilingual pre- trained model,
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y . Xu, W. Zheng, X. Xia, W. L. Tam, Z. Ma, Y . Xue, J. Zhai, W. Chen, P. Zhang, Y . Dong, and J. Tang, “Glm-130b: An open bilingual pre- trained model,” 2023
2023
-
[152]
Flm-101b: An open llm and how to train it with $100 k budget,
X. Li, Y . Yao, X. Jiang, X. Fang, X. Meng, S. Fan, P. Han, J. Li, L. Du, B. Qin et al., “Flm-101b: An open llm and how to train it with $100 k budget,” arXiv preprint arXiv:2309.03852 , 2023
2023 arXiv
-
[153]
Flamingo: a visual language model for few-shot learning,
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems , vol. 35, pp. 23 716–23 736, 2022
2022
-
[154]
Gpt-4 technical report,
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023 arXiv
-
[155]
Minigpt-4: En- hancing vision-language understanding with advanced large language models,
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: En- hancing vision-language understanding with advanced large language models,” arXiv preprint arXiv:2304.10592 , 2023
2023 arXiv
-
[156]
mplug-owl: Modularization empowers large language models with multimodality,
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y . Zhou, J. Wang, A. Hu, P. Shi, Y . Shi et al. , “mplug-owl: Modularization empowers large language models with multimodality,” arXiv preprint arXiv:2304.14178 , 2023
2023 arXiv
-
[157]
Cogvlm: Visual expert for pretrained language models,
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y . Wang, J. Ji, Z. Yang, L. Zhao, X. Song et al., “Cogvlm: Visual expert for pretrained language models,” arXiv preprint arXiv:2311.03079 , 2023
2023 arXiv
-
[158]
Imagebind: One embedding space to bind them all,
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 15 180–15 190
2023
-
[159]
Pandagpt: One model to instruction-follow them all,
Y . Su, T. Lan, H. Li, J. Xu, Y . Wang, and D. Cai, “Pandagpt: One model to instruction-follow them all,” arXiv preprint arXiv:2305.16355, 2023
2023 arXiv
-
[160]
Next-gpt: Any-to-any multimodal llm,
S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua, “Next-gpt: Any-to-any multimodal llm,” arXiv preprint arXiv:2309.05519 , 2023
2023 arXiv
-
[161]
Onellm: One framework to align all modalities with language,
J. Han, K. Gong, Y . Zhang, J. Wang, K. Zhang, D. Lin, Y . Qiao, P. Gao, and X. Yue, “Onellm: One framework to align all modalities with language,” arXiv preprint arXiv:2312.03700 , 2023
2023 arXiv
-
[162]
Gemini: a family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, Y . Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth et al., “Gemini: a family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805 , 2023
2023 arXiv
-
[163]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser et al. , “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” arXiv preprint arXiv:2403.05530 , 2024
2024 arXiv
-
[164]
Shikra: Unleashing multimodal llm’s referential dialogue magic,
K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao, “Shikra: Unleashing multimodal llm’s referential dialogue magic,” arXiv preprint arXiv:2306.15195 , 2023
2023 arXiv
-
[165]
Ferret: Refer and ground anything anywhere at any granularity,
H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S.-F. Chang, and Y . Yang, “Ferret: Refer and ground anything anywhere at any granularity,” arXiv preprint arXiv:2310.07704 , 2023
2023 arXiv
-
[166]
Pix2seq: A language modeling framework for object detection,
T. Chen, S. Saxena, L. Li, D. J. Fleet, and G. Hinton, “Pix2seq: A language modeling framework for object detection,” 2022
2022
-
[167]
Next- chat: An lmm for chat, detection and segmentation,
A. Zhang, L. Zhao, C.-W. Xie, Y . Zheng, W. Ji, and T.-S. Chua, “Next- chat: An lmm for chat, detection and segmentation,” arXiv preprint arXiv:2311.04498, 2023
2023 arXiv
-
[168]
Pangu- {\Sigma}: To- wards trillion parameter language model with sparse heterogeneous computing,
X. Ren, P. Zhou, X. Meng, X. Huang, Y . Wang, W. Wang, P. Li, X. Zhang, A. Podolskiy, G. Arshinov et al. , “Pangu- {\Sigma}: To- wards trillion parameter language model with sparse heterogeneous computing,” arXiv preprint arXiv:2303.10845 , 2023
2023 arXiv
-
[169]
Mistral 7b,
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier et al., “Mistral 7b,” arXiv preprint arXiv:2310.06825 , 2023
2023 arXiv
-
[170]
Mixtral of experts,
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bam- ford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand et al. , “Mixtral of experts,” arXiv preprint arXiv:2401.04088 , 2024. MANUSCRIPT 38
2024 arXiv
-
[171]
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, 2022
2022
-
[172]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” 2019
2019
-
[173]
Tinybert: Distilling bert for natural language understanding,
X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, “Tinybert: Distilling bert for natural language understanding,” arXiv preprint arXiv:1909.10351 , 2019
1909 arXiv
-
[174]
Rethinking optimization and architecture for tiny language models,
Y . Tang, F. Liu, Y . Ni, Y . Tian, Z. Bai, Y .-Q. Hu, S. Liu, S. Jui, K. Han, and Y . Wang, “Rethinking optimization and architecture for tiny language models,” arXiv preprint arXiv:2402.02791 , 2024
2024 arXiv
-
[175]
Tinyllama: An open-source small language model,
P. Zhang, G. Zeng, T. Wang, and W. Lu, “Tinyllama: An open-source small language model,” arXiv preprint arXiv:2401.02385 , 2024
2024 arXiv
-
[176]
Phi- 3 technical report: A highly capable language model locally on your phone,
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl et al. , “Phi- 3 technical report: A highly capable language model locally on your phone,” arXiv preprint arXiv:2404.14219 , 2024
2024 arXiv
-
[177]
A simple and effec- tive pruning approach for large language models,
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effec- tive pruning approach for large language models,” arXiv preprint arXiv:2306.11695, 2023
2023 arXiv
-
[178]
Llm-pruner: On the structural pruning of large language models,
X. Ma, G. Fang, and X. Wang, “Llm-pruner: On the structural pruning of large language models,” Advances in neural information processing systems, vol. 36, 2024
2024
-
[179]
Loraprune: Pruning meets low-rank parameter-efficient fine-tuning,
M. Zhang, H. Chen, C. Shen, Z. Yang, L. Ou, X. Yu, and B. Zhuang, “Loraprune: Pruning meets low-rank parameter-efficient fine-tuning,” 2023
2023
-
[180]
Lorashear: Efficient large language model structured pruning and knowledge recovery,
T. Chen, T. Ding, B. Yadav, I. Zharkov, and L. Liang, “Lorashear: Efficient large language model structured pruning and knowledge recovery,” arXiv preprint arXiv:2310.18356 , 2023
2023 arXiv
-
[181]
Fluctuation-based adaptive structured pruning for large language models,
Y . An, X. Zhao, T. Yu, M. Tang, and J. Wang, “Fluctuation-based adaptive structured pruning for large language models,” arXiv preprint arXiv:2312.11983, 2023
2023 arXiv
-
[182]
One-shot sensitivity-aware mixed sparsity pruning for large language models,
H. Shao, B. Liu, and Y . Qian, “One-shot sensitivity-aware mixed sparsity pruning for large language models,” arXiv preprint arXiv:2310.09499, 2023
2023 arXiv
-
[183]
Compresso: Structured pruning with collaborative prompting learns compact large language models,
S. Guo, J. Xu, L. L. Zhang, and M. Yang, “Compresso: Structured pruning with collaborative prompting learns compact large language models,” arXiv preprint arXiv:2310.05015 , 2023
2023 arXiv
-
[184]
Sheared llama: Accelerating language model pre-training via structured pruning,
M. Xia, T. Gao, Z. Zeng, and D. Chen, “Sheared llama: Accelerating language model pre-training via structured pruning,” arXiv preprint arXiv:2310.06694, 2023
2023 arXiv
-
[185]
Pruning large language models via accuracy predictor,
Y . Ji, Y . Cao, and J. Liu, “Pruning large language models via accuracy predictor,” arXiv preprint arXiv:2309.09507 , 2023
2023 arXiv
-
[186]
Beyond size: How gradients shape pruning decisions in large language models,
R. J. Das, L. Ma, and Z. Shen, “Beyond size: How gradients shape pruning decisions in large language models,” arXiv preprint arXiv:2311.04902, 2023
2023 arXiv
-
[187]
Dynamic context pruning for efficient and interpretable autoregressive transformers,
S. Anagnostidis, D. Pavllo, L. Biggio, L. Noci, A. Lucchi, and T. Hofmann, “Dynamic context pruning for efficient and interpretable autoregressive transformers,” Advances in Neural Information Process- ing Systems, vol. 36, 2024
2024
-
[188]
ZipLM: Inference-Aware Struc- tured Pruning of Language Models,
E. Kurti ´c, E. Frantar, and D. Alistarh, “ZipLM: Inference-Aware Struc- tured Pruning of Language Models,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[189]
DPHuBERT: Joint distillation and pruning of self-supervised speech models,
Y . Peng, Y . Sudo, S. Muhammad, and S. Watanabe, “DPHuBERT: Joint distillation and pruning of self-supervised speech models,” arXiv preprint arXiv:2305.17651, 2023
2023 arXiv
-
[190]
Structured Prun- ing of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding,
Y . Peng, K. Kim, F. Wu, P. Sridhar, and S. Watanabe, “Structured Prun- ing of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
-
[191]
MoPE-CLIP: Structured Pruning for Efficient Vision- Language Models with Module-wise Pruning Error Metric,
H. Lin, H. Bai, Z. Liu, L. Hou, M. Sun, L. Song, Y . Wei, and Z. Sun, “MoPE-CLIP: Structured Pruning for Efficient Vision- Language Models with Module-wise Pruning Error Metric,” arXiv preprint arXiv:2403.07839, 2024
2024 arXiv
-
[192]
X-pruner: explainable pruning for vision trans- formers,
L. Yu and W. Xiang, “X-pruner: explainable pruning for vision trans- formers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 24 355–24 363
2023
-
[193]
Structural pruning for diffusion models,
G. Fang, X. Ma, and X. Wang, “Structural pruning for diffusion models,” Advances in neural information processing systems , vol. 36, 2024
2024
-
[194]
A unified pruning framework for vision transform- ers,
H. Yu and J. Wu, “A unified pruning framework for vision transform- ers,” Science China Information Sciences , vol. 66, no. 7, p. 179101, 2023
2023
-
[195]
CAP: Correlation-Aware Pruning for Highly-Accurate Sparse Vision Mod- els,
D. Kuznedelev, E. Kurti ´c, E. Frantar, and D. Alistarh, “CAP: Correlation-Aware Pruning for Highly-Accurate Sparse Vision Mod- els,” in Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curr...
2023
-
[196]
Pada: Pruning assisted domain adaptation for self-supervised speech representations,
V . S. Lodagala, S. Ghosh, and S. Umesh, “Pada: Pruning assisted domain adaptation for self-supervised speech representations,” in 2022 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2023, pp. 136–143
2022
-
[197]
Smoothquant: Accurate and efficient post-training quantization for large language models,
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning. PMLR, 2023, pp. 38 087–38 099
2023
-
[198]
Rptq: Reorder-based post-training quantization for large language models,
Z. Yuan, L. Niu, J. Liu, W. Liu, X. Wang, Y . Shang, G. Sun, Q. Wu, J. Wu, and B. Wu, “Rptq: Reorder-based post-training quantization for large language models,” arXiv preprint arXiv:2304.01089 , 2023
2023 arXiv
-
[199]
Loftq: Lora-fine-tuning-aware quantization for large language models,
Y . Li, Y . Yu, C. Liang, P. He, N. Karampatziakis, W. Chen, and T. Zhao, “Loftq: Lora-fine-tuning-aware quantization for large language models,” arXiv preprint arXiv:2310.08659 , 2023
2023 arXiv
-
[200]
Outlier suppression+: Accurate quantization of large language mod- els by equivalent and optimal shifting and scaling,
X. Wei, Y . Zhang, Y . Li, X. Zhang, R. Gong, J. Guo, and X. Liu, “Outlier suppression+: Accurate quantization of large language mod- els by equivalent and optimal shifting and scaling,” arXiv preprint arXiv:2304.09145, 2023
2023 arXiv
-
[201]
FPTQ: Fine-grained Post-Training Quantization for Large Language Models,
Q. Li, Y . Zhang, L. Li, P. Yao, B. Zhang, X. Chu, Y . Sun, L. Du, and Y . Xie, “FPTQ: Fine-grained Post-Training Quantization for Large Language Models,” arXiv preprint arXiv:2308.15987 , 2023
2023 arXiv
-
[202]
Owq: Lessons learned from activation outliers for weight quantization in large language models,
C. Lee, J. Jin, T. Kim, H. Kim, and E. Park, “Owq: Lessons learned from activation outliers for weight quantization in large language models,” arXiv preprint arXiv:2306.02272 , 2023
2023 arXiv
-
[203]
Awq: Activation-aware weight quantization for llm compression and accel- eration,
J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, and S. Han, “Awq: Activation-aware weight quantization for llm compression and accel- eration,” arXiv preprint arXiv:2306.00978 , 2023
2023 arXiv
-
[204]
Integer or floating point? new outlooks for low-bit quantization on large language models,
Y . Zhang, L. Zhao, S. Cao, W. Wang, T. Cao, F. Yang, M. Yang, S. Zhang, and N. Xu, “Integer or floating point? new outlooks for low-bit quantization on large language models,” arXiv preprint arXiv:2305.12356, 2023
2023 arXiv
-
[205]
Omniquant: Omnidirectionally calibrated quan- tization for large language models,
W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, K. Zhang, P. Gao, Y . Qiao, and P. Luo, “Omniquant: Omnidirectionally calibrated quan- tization for large language models,” arXiv preprint arXiv:2308.13137 , 2023
2023 arXiv
-
[206]
IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact,
R. Liu, H. Bai, H. Lin, Y . Li, H. Gao, Z. Xu, L. Hou, J. Yao, and C. Yuan, “IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact,” arXiv preprint arXiv:2403.01241, 2024
2024 arXiv
-
[207]
Memory-Efficient Fine-Tuning of Compressed Large Language Mod- els via sub-4-bit Integer Quantization,
J. Kim, J. H. Lee, S. Kim, J. Park, K. M. Yoo, S. J. Kwon, and D. Lee, “Memory-Efficient Fine-Tuning of Compressed Large Language Mod- els via sub-4-bit Integer Quantization,” in Advances in Neural Informa- tion Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M...
2023
-
[208]
Qllm: Accurate and efficient low-bitwidth quantization for large language models,
J. Liu, R. Gong, X. Wei, Z. Dong, J. Cai, and B. Zhuang, “Qllm: Accurate and efficient low-bitwidth quantization for large language models,” arXiv preprint arXiv:2310.08041 , 2023
2023 arXiv
-
[209]
Llm-qat: Data-free quantization aware training for large language models,
Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y . Mehdad, Y . Shi, R. Kr- ishnamoorthi, and V . Chandra, “Llm-qat: Data-free quantization aware training for large language models,” arXiv preprint arXiv:2305.17888, 2023
2023 arXiv
-
[210]
Quip: 2-bit quanti- zation of large language models with guarantees,
J. Chee, Y . Cai, V . Kuleshov, and C. M. De Sa, “Quip: 2-bit quanti- zation of large language models with guarantees,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[211]
Norm tweaking: High- performance low-bit quantization of large language models,
L. Li, Q. Li, B. Zhang, and X. Chu, “Norm tweaking: High- performance low-bit quantization of large language models,” arXiv preprint arXiv:2309.02784, 2023
2023 arXiv
-
[212]
Zeroquant-v2: Exploring post-training quantization in llms from comprehensive study to low rank compensation,
Z. Yao, X. Wu, C. Li, S. Youn, and Y . He, “Zeroquant-v2: Exploring post-training quantization in llms from comprehensive study to low rank compensation,” arXiv preprint arXiv:2303.08302 , 2023
2023 arXiv
-
[213]
Qa-lora: Quantization-aware low-rank adaptation of large language models,
Y . Xu, L. Xie, X. Gu, X. Chen, H. Chang, H. Zhang, Z. Chen, X. Zhang, and Q. Tian, “Qa-lora: Quantization-aware low-rank adaptation of large language models,” arXiv preprint arXiv:2309.14717 , 2023
2023 arXiv
-
[214]
Int2.1: Towards fine-tunable quantized large language models with error cor- rection through low-rank adaptation,
Y . Chai, J. Gkountouras, G. G. Ko, D. Brooks, and G.-Y . Wei, “Int2.1: Towards fine-tunable quantized large language models with error cor- rection through low-rank adaptation,”arXiv preprint arXiv:2306.08162, 2023
2023 arXiv
-
[215]
The Quantization Model of Neural Scaling,
E. J. Michaud, Z. Liu, U. Girit, and M. Tegmark, “The Quantization Model of Neural Scaling,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
-
[216]
Quantized transformer language model implementations on edge devices,
M. W. U. Rahman, M. M. Abrar, H. G. Copening, S. Hariri, S. Shao, P. Satam, and S. Salehi, “Quantized transformer language model implementations on edge devices,” arXiv preprint arXiv:2310.03971 , 2023. MANUSCRIPT 39
2023 arXiv
-
[217]
Distilling step-by-step! outper- forming larger language models with less training data and smaller model sizes,
C.-Y . Hsieh, C.-L. Li, C.-K. Yeh, H. Nakhost, Y . Fujii, A. Ratner, R. Krishna, C.-Y . Lee, and T. Pfister, “Distilling step-by-step! outper- forming larger language models with less training data and smaller model sizes,” arXiv preprint arXiv:2305.02301 , 2023
2023 arXiv
-
[218]
Zephyr: Direct distillation of lm alignment,
L. Tunstall, E. Beeching, N. Lambert, N. Rajani, K. Rasul, Y . Belkada, S. Huang, L. von Werra, C. Fourrier, N. Habib et al., “Zephyr: Direct distillation of lm alignment,” arXiv preprint arXiv:2310.16944 , 2023
2023 arXiv
-
[219]
Lion: Adversarial distillation of closed-source large language model,
Y . Jiang, C. Chan, M. Chen, and W. Wang, “Lion: Adversarial distillation of closed-source large language model,” arXiv preprint arXiv:2305.12870, 2023
2023 arXiv
-
[220]
Pad: Program- aided distillation specializes large models in reasoning,
X. Zhu, B. Qi, K. Zhang, X. Long, and B. Zhou, “Pad: Program- aided distillation specializes large models in reasoning,” arXiv preprint arXiv:2305.13888, 2023
2023 arXiv
-
[221]
Distillspec: Improving speculative decoding via knowledge distillation,
Y . Zhou, K. Lyu, A. S. Rawat, A. K. Menon, A. Rostamizadeh, S. Ku- mar, J.-F. Kagy, and R. Agarwal, “Distillspec: Improving speculative decoding via knowledge distillation,” arXiv preprint arXiv:2310.08461, 2023
2023 arXiv
-
[222]
Knowledge distillation of llm for education,
E. Latif, L. Fang, P. Ma, and X. Zhai, “Knowledge distillation of llm for education,” arXiv preprint arXiv:2312.15842 , 2023
2023 arXiv
-
[223]
Less is more: Task-aware layer-wise distillation for language model compres- sion,
C. Liang, S. Zuo, Q. Zhang, P. He, W. Chen, and T. Zhao, “Less is more: Task-aware layer-wise distillation for language model compres- sion,” in International Conference on Machine Learning . PMLR, 2023, pp. 20 852–20 867
2023
-
[224]
Homodistil: Homotopic task-agnostic distillation of pre-trained transformers,
C. Liang, H. Jiang, Z. Li, X. Tang, B. Yin, and T. Zhao, “Homodistil: Homotopic task-agnostic distillation of pre-trained transformers,” arXiv preprint arXiv:2302.09632, 2023
2023 arXiv
-
[225]
Symbolic chain-of-thought distillation: Small models can also
L. H. Li, J. Hessel, Y . Yu, X. Ren, K.-W. Chang, and Y . Choi, “Symbolic chain-of-thought distillation: Small models can also” think” step-by-step,” arXiv preprint arXiv:2306.14050 , 2023
2023 arXiv
-
[226]
Evolving Knowledge Distillation with Large Language Models and Active Learning,
C. Liu, Y . Kang, F. Zhao, K. Kuang, Z. Jiang, C. Sun, and F. Wu, “Evolving Knowledge Distillation with Large Language Models and Active Learning,” arXiv preprint arXiv:2403.06414 , 2024
2024 arXiv
-
[227]
Minimal Distillation Schedule for Extreme Language Model Com- pression,
C. Zhang, Y . Yang, Q. Wang, J. Liu, J. Wang, W. Wu, and D. Song, “Minimal Distillation Schedule for Extreme Language Model Com- pression,” in Findings of the Association for Computational Linguistics: EACL 2024, 2024, pp. 1378–1394
2024
-
[228]
Slam: Student-label mixing for distillation with unlabeled ex- amples,
V . Kontonis, F. Iliopoulos, K. Trinh, C. Baykal, G. Menghani, and E. Vee, “Slam: Student-label mixing for distillation with unlabeled ex- amples,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[229]
Universalner: Targeted distillation from large language models for open named entity recognition,
W. Zhou, S. Zhang, Y . Gu, M. Chen, and H. Poon, “Universalner: Targeted distillation from large language models for open named entity recognition,” arXiv preprint arXiv:2308.03279 , 2023
2023 arXiv
-
[230]
Scott: Self-consistent chain-of-thought distillation,
P. Wang, Z. Wang, Z. Li, Y . Gao, B. Yin, and X. Ren, “Scott: Self-consistent chain-of-thought distillation,” arXiv preprint arXiv:2305.01879, 2023
2023 arXiv
-
[231]
Detrdistill: A universal knowledge distillation framework for detr- families,
J. Chang, S. Wang, H.-M. Xu, Z. Chen, C. Yang, and F. Zhao, “Detrdistill: A universal knowledge distillation framework for detr- families,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 6898–6908
2023
-
[232]
Boosting accuracy and robustness of student models via adaptive adversarial distillation,
B. Huang, M. Chen, Y . Wang, J. Lu, M. Cheng, and W. Wang, “Boosting accuracy and robustness of student models via adaptive adversarial distillation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 24 668–24 677
2023
-
[233]
Dime-fm: Distilling multimodal and efficient foundation models,
X. Sun, P. Zhang, P. Zhang, H. Shah, K. Saenko, and X. Xia, “Dime-fm: Distilling multimodal and efficient foundation models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 15 521–15 533
2023
-
[234]
Contextualization distillation from large language model for knowledge graph completion,
D. Li, Z. Tan, T. Chen, and H. Liu, “Contextualization distillation from large language model for knowledge graph completion,” arXiv preprint arXiv:2402.01729, 2024
2024 arXiv
-
[235]
Large Language Model Meets Graph Neural Network in Knowledge Distillation,
S. Hu, G. Zou, S. Yang, B. Zhang, and Y . Chen, “Large Language Model Meets Graph Neural Network in Knowledge Distillation,” arXiv preprint arXiv:2402.05894, 2024
2024 arXiv
-
[236]
On Good Practices for Task-Specific Distillation of Large Pretrained Models,
J. Marrie, M. Arbel, J. Mairal, and D. Larlus, “On Good Practices for Task-Specific Distillation of Large Pretrained Models,” arXiv preprint arXiv:2402.11305, 2024
2024 arXiv
-
[237]
Dafkd: Domain-aware federated knowledge distillation,
H. Wang, Y . Li, W. Xu, R. Li, Y . Zhan, and Z. Zeng, “Dafkd: Domain-aware federated knowledge distillation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2023, pp. 20 412–20 421
2023
-
[238]
Edgeadaptor: Online configuration adaption, model selection and resource provisioning for edge dnn inference serving at scale,
K. Zhao, Z. Zhou, X. Chen, R. Zhou, X. Zhang, S. Yu, and D. Wu, “Edgeadaptor: Online configuration adaption, model selection and resource provisioning for edge dnn inference serving at scale,” IEEE Transactions on Mobile Computing , 2022
2022
-
[239]
Sti: Turbocharge nlp inference at the edge via elastic pipelining,
L. Guo, W. Choe, and F. X. Lin, “Sti: Turbocharge nlp inference at the edge via elastic pipelining,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2023, pp. 791–803
2023
-
[240]
Tabi: An efficient multi- level inference system for large language models,
Y . Wang, K. Chen, H. Tan, and K. Guo, “Tabi: An efficient multi- level inference system for large language models,” in Proceedings of the Eighteenth European Conference on Computer Systems , 2023, pp. 233–248
2023
-
[241]
Consistent Acceler- ated Inference via Confident Adaptive Transformers,
T. Schuster, A. Fisch, T. Jaakkola, and R. Barzilay, “Consistent Acceler- ated Inference via Confident Adaptive Transformers,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 4962–4979
2021
-
[242]
Virtual reality solutions employing artificial intelligence methods: A systematic literature review,
T. Ribeiro de Oliveira, B. Biancardi Rodrigues, M. Moura da Silva, R. Antonio N. Spinass ´e, G. Giesen Ludke, M. Ruy Soares Gaudio, G. Iglesias Rocha Gomes, L. Guio Cotini, D. da Silva Vargens, M. Queiroz Schimidt et al. , “Virtual reality solutions employing artificial intell...
2023
-
[243]
Use of AI-based tools for healthcare purposes: a survey study from consumers’ perspectives,
P. Esmaeilzadeh, “Use of AI-based tools for healthcare purposes: a survey study from consumers’ perspectives,” BMC medical informatics and decision making , vol. 20, pp. 1–19, 2020
2020
-
[244]
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding,
Z. Chen, A. May, R. Svirschevski, Y . Huang, M. Ryabinin, Z. Jia, and B. Chen, “Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding,” arXiv preprint arXiv:2402.12374 , 2024
2024 arXiv
-
[245]
Minions: Accelerating Large Language Model Inference with Adaptive and Collective Speculative Decoding,
S. Wang, H. Yang, X. Wang, T. Liu, P. Wang, X. Liang, K. Ma, T. Feng, X. You, Y . Baoet al., “Minions: Accelerating Large Language Model Inference with Adaptive and Collective Speculative Decoding,” arXiv preprint arXiv:2402.15678, 2024
2024 arXiv
-
[246]
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs,
S. Ge, Y . Zhang, L. Liu, M. Zhang, J. Han, and J. Gao, “Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs,” in The Twelfth International Conference on Learning Representations , 2023
2023
-
[247]
Efficient streaming language models with attention sinks,
G. Xiao, Y . Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in The Twelfth International Conference on Learning Representations , 2023
2023
-
[248]
H2o: Heavy-hitter oracle for efficient generative inference of large language models,
Z. Zhang, Y . Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y . Tian, C. R´e, C. Barrett et al., “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[249]
Constraint-aware and ranking-distilled token pruning for efficient transformer inference,
J. Li, L. L. Zhang, J. Xu, Y . Wang, S. Yan, Y . Xia, Y . Yang, T. Cao, H. Sun, W. Deng et al., “Constraint-aware and ranking-distilled token pruning for efficient transformer inference,” in Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining ,...
2023
-
[250]
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models,
H. Jiang, Q. Wu, C.-Y . Lin, Y . Yang, and L. Qiu, “LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 13 358–13 376
2023
-
[251]
Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression,
H. Jiang, Q. Wu, X. Luo, D. Li, C.-Y . Lin, Y . Yang, and L. Qiu, “Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression,” arXiv preprint arXiv:2310.06839 , 2023
2023 arXiv
-
[252]
Compressing Context to En- hance Inference Efficiency of Large Language Models,
Y . Li, B. Dong, F. Guerin, and C. Lin, “Compressing Context to En- hance Inference Efficiency of Large Language Models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 6342–6353
2023
-
[253]
Learned token pruning for transformers,
S. Kim, S. Shen, D. Thorsley, A. Gholami, W. Kwon, J. Hassoun, and K. Keutzer, “Learned token pruning for transformers,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022, pp. 784–794
2022
-
[254]
AdapLeR: Speeding up Inference by Adaptive Length Reduction,
A. Modarressi, H. Mohebbi, and M. T. Pilehvar, “AdapLeR: Speeding up Inference by Adaptive Length Reduction,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1–15
2022
-
[255]
Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search,
G. Kim and K. Cho, “Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume...
2021
-
[256]
Power-bert: Accelerating bert inference via progressive word-vector elimination,
S. Goyal, A. R. Choudhury, S. Raje, V . Chakaravarthy, Y . Sabharwal, and A. Verma, “Power-bert: Accelerating bert inference via progressive word-vector elimination,” in International Conference on Machine Learning. PMLR, 2020, pp. 3690–3699
2020
-
[257]
In-context Autoencoder for Context Compression in a Large Language Model,
T. Ge, H. Jing, L. Wang, X. Wang, S.-Q. Chen, and F. Wei, “In-context Autoencoder for Context Compression in a Large Language Model,” in The Twelfth International Conference on Learning Representations , 2023. MANUSCRIPT 40
2023
-
[258]
Learning to compress prompts with gist tokens,
J. Mu, X. Li, and N. Goodman, “Learning to compress prompts with gist tokens,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[259]
Adapting Language Models to Compress Contexts,
A. Chevalier, A. Wettig, A. Ajith, and D. Chen, “Adapting Language Models to Compress Contexts,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 3829–3846
2023
-
[260]
SYNERGISTIC PATCH PRUN- ING FOR VISION TRANS-FORMER: UNIFYING INTRA-&INTER- LAYER PATCH IMPORTANCE
Y . Zhang, L. Wei, and N. M. Freris, “SYNERGISTIC PATCH PRUN- ING FOR VISION TRANS-FORMER: UNIFYING INTRA-&INTER- LAYER PATCH IMPORTANCE.”
-
[261]
A Simple Romance Between Multi-Exit Vision Transformer and Token Reduction,
D. Liu, M. Kan, S. Shan, and C. Xilin, “A Simple Romance Between Multi-Exit Vision Transformer and Token Reduction,” in The Twelfth International Conference on Learning Representations , 2023
2023
-
[262]
Heatvit: Hardware-efficient adaptive token prun- ing for vision transformers,
P. Dong, M. Sun, A. Lu, Y . Xie, K. Liu, Z. Kong, X. Meng, Z. Li, X. Lin, Z. Fang et al., “Heatvit: Hardware-efficient adaptive token prun- ing for vision transformers,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2023, pp. 442–455
2023
-
[263]
Patch slimming for efficient vision transformers,
Y . Tang, K. Han, Y . Wang, C. Xu, J. Guo, C. Xu, and D. Tao, “Patch slimming for efficient vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 165–12 174
2022
-
[264]
Dynamicvit: Efficient vision transformers with dynamic token sparsification,
Y . Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “Dynamicvit: Efficient vision transformers with dynamic token sparsification,” Ad- vances in neural information processing systems , vol. 34, pp. 13 937– 13 949, 2021
2021
-
[265]
Token Merging: Your ViT But Faster,
D. Bolya, C.-Y . Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token Merging: Your ViT But Faster,” in The Eleventh International Conference on Learning Representations , 2022
2022
-
[266]
EViT: Expediting Vision Transformers via Token Reorganizations,
Y . Liang, G. Chongjian, Z. Tong, Y . Song, J. Wang, and P. Xie, “EViT: Expediting Vision Transformers via Token Reorganizations,” in International Conference on Learning Representations , 2021
2021
-
[267]
DiffRate: Differentiable Compression Rate for Efficient Vision Transformers,
M. Chen, W. Shao, P. Xu, M. Lin, K. Zhang, F. Chao, R. Ji, Y . Qiao, and P. Luo, “DiffRate: Differentiable Compression Rate for Efficient Vision Transformers,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, 2023, pp. 17 118–17 128
2023
-
[268]
Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision Transformers,
S. Long, Z. Zhao, J. Pi, S. Wang, and J. Wang, “Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision Transformers,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2023, pp. 10 334– 10 343
2023
-
[269]
Joint token pruning and squeezing towards more aggressive compression of vision trans- formers,
S. Wei, T. Ye, S. Zhang, Y . Tang, and J. Liang, “Joint token pruning and squeezing towards more aggressive compression of vision trans- formers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2092–2101
2023
-
[270]
PuMer: Pruning and Merging Tokens for Efficient Vision Language Models,
Q. Cao, B. Paranjape, and H. Hajishirzi, “PuMer: Pruning and Merging Tokens for Efficient Vision Language Models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12 890–12 903
2023
-
[271]
TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding,
S. Ren, S. Chen, S. Li, X. Sun, and L. Hou, “TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp. 932–947
2023
-
[272]
Prune spatio-temporal tokens by semantic-aware temporal accumulation,
S. Ding, P. Zhao, X. Zhang, R. Qian, H. Xiong, and Q. Tian, “Prune spatio-temporal tokens by semantic-aware temporal accumulation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 16 945–16 956
2023
-
[273]
Efficient video transformers with spatial-temporal token selection,
J. Wang, X. Yang, H. Li, L. Liu, Z. Wu, and Y .-G. Jiang, “Efficient video transformers with spatial-temporal token selection,” in European Conference on Computer Vision . Springer, 2022, pp. 69–86
2022
-
[274]
Efficient Transform- ers: A Survey,
Y . Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient Transform- ers: A Survey,” ACM Comput. Surv., vol. 55, no. 6, dec 2022
2022
-
[275]
Which tokens to use? investigating token reduction in vision transformers,
J. B. Haurum, S. Escalera, G. W. Taylor, and T. B. Moeslund, “Which tokens to use? investigating token reduction in vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 773–783
2023
-
[276]
OTAS: An Elastic Transformer Serving System via Token Adapta- tion,
J. Chen, W. Xu, Z. Hong, S. Guo, H. Wang, J. Zhang, and D. Zeng, “OTAS: An Elastic Transformer Serving System via Token Adapta- tion,” 2024
2024
-
[277]
Adaptive computation with elastic input sequence,
F. Xue, V . Likhosherstov, A. Arnab, N. Houlsby, M. Dehghani, and Y . You, “Adaptive computation with elastic input sequence,” in Inter- national Conference on Machine Learning. PMLR, 2023, pp. 38 971– 38 988
2023
-
[278]
Levels of AGI: Operationalizing Progress on the Path to AGI,
M. R. Morris, J. Sohl-dickstein, N. Fiedel, T. Warkentin, A. Dafoe, A. Faust, C. Farabet, and S. Legg, “Levels of AGI: Operationalizing Progress on the Path to AGI,” arXiv preprint arXiv:2311.02462, 2023
2023
-
[279]
AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors,
W. Chen, Y . Su, J. Zuo, C. Yang, C. Yuan, C.-M. Chan, yang Yu, Y . Lu, Y .-H. Hung, C. Qian, Y . Qin, X. Cong, R. Xie, Z. Liu, M. Sun, and J. Zhou, “AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors,” in The Twelfth International Conference o...
2024
-
[280]
Agentbench: Evaluating llms as agents,
X. Liu, H. Yu, H. Zhang, Y . Xu, X. Lei, H. Lai, Y . Gu, H. Ding, K. Men, K. Yang et al., “Agentbench: Evaluating llms as agents,” arXiv preprint arXiv:2308.03688, 2023
2023 arXiv
-
[281]
Building cooperative embodied agents modularly with large language models,
H. Zhang, W. Du, J. Shan, Q. Zhou, Y . Du, J. B. Tenenbaum, T. Shu, and C. Gan, “Building cooperative embodied agents modularly with large language models,” arXiv preprint arXiv:2307.02485 , 2023
2023 arXiv
-
[282]
Evaluating multi-agent coordination abilities in large language models,
S. Agashe, Y . Fan, and X. E. Wang, “Evaluating multi-agent coordination abilities in large language models,” arXiv preprint arXiv:2310.03903, 2023
2023 arXiv
-
[283]
Grounded Decod- ing: Guiding Text Generation with Grounded Models for Embodied Agents,
W. Huang, F. Xia, D. Shah, D. Driess, A. Zeng, Y . Lu, P. Florence, I. Mordatch, S. Levine, K. Hausman, and B. Ichter, “Grounded Decod- ing: Guiding Text Generation with Grounded Models for Embodied Agents,” in Conference on Neural Information Processing Systems , 2023
2023
-
[284]
Empowering Conversational Agents using Semantic In-Context Learning,
A. Omidvar and A. An, “Empowering Conversational Agents using Semantic In-Context Learning,” in Annual Meeting of the Association for Computational Linguistics , 2023
2023
-
[285]
Multi-agent collaboration: Harnessing the power of intelligent llm agents,
Y . Talebirad and A. Nadiri, “Multi-agent collaboration: Harnessing the power of intelligent llm agents,” arXiv preprint arXiv:2306.03314, 2023
2023 arXiv
-
[286]
Dynamic llm-agent network: An llm-agent collaboration framework with agent team opti- mization,
Z. Liu, Y . Zhang, P. Li, Y . Liu, and D. Yang, “Dynamic llm-agent network: An llm-agent collaboration framework with agent team opti- mization,” arXiv preprint arXiv:2310.02170 , 2023
2023 arXiv
-
[287]
Mindagent: Emergent gaming interaction,
R. Gong, Q. Huang, X. Ma, H. V o, Z. Durante, Y . Noda, Z. Zheng, S.-C. Zhu, D. Terzopoulos, L. Fei-Fei et al. , “Mindagent: Emergent gaming interaction,” arXiv preprint arXiv:2309.09971 , 2023
2023 arXiv
-
[288]
Describe, explain, plan and select: interactive planning with LLMs enables open- world multi-task agents,
Z. Wang, S. Cai, G. Chen, A. Liu, X. S. Ma, and Y . Liang, “Describe, explain, plan and select: interactive planning with LLMs enables open- world multi-task agents,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[289]
Large language models as common- sense knowledge for large-scale task planning,
Z. Zhao, W. S. Lee, and D. Hsu, “Large language models as common- sense knowledge for large-scale task planning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[290]
Large language models can implement policy iteration,
E. Brooks, L. Walls, R. L. Lewis, and S. Singh, “Large language models can implement policy iteration,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[291]
Large language models of code fail at completing code with potential bugs,
T. Dinh, J. Zhao, S. Tan, R. Negrinho, L. Lausen, S. Zha, and G. Karypis, “Large language models of code fail at completing code with potential bugs,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[292]
Lever- aging pre-trained large language models to construct and utilize world models for model-based task planning,
L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati, “Lever- aging pre-trained large language models to construct and utilize world models for model-based task planning,” Advances in Neural Informa- tion Processing Systems , vol. 36, 2024
2024
-
[293]
Avalon’s Game of Thoughts: Battle Against Deception through Recursive Contemplation,
S. Wang, C. Liu, Z. Zheng, S. Qi, S. Chen, Q. Yang, A. Zhao, C. Wang, S. Song, and G. Huang, “Avalon’s Game of Thoughts: Battle Against Deception through Recursive Contemplation,” arXiv preprint arXiv:2310.01320, 2023
2023 arXiv
-
[294]
Prompted LLMs as Chatbot Modules for Long Open-domain Conversation,
G. Lee, V . Hartmann, J. Park, D. Papailiopoulos, and K. Lee, “Prompted LLMs as Chatbot Modules for Long Open-domain Conversation,” in Annual Meeting of the Association for Computational Linguistics , 2023
2023
-
[295]
Large Lan- guage Models Are Semi-Parametric Reinforcement Learning Agents,
D. Zhang, L. Chen, S. Zhang, H. Xu, Z. Zhao, and K. Yu, “Large Lan- guage Models Are Semi-Parametric Reinforcement Learning Agents,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[296]
Augmenting language models with long-term memory,
W. Wang, L. Dong, H. Cheng, X. Liu, X. Yan, J. Gao, and F. Wei, “Augmenting language models with long-term memory,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[297]
Memory-augmented llm personalization with short-and long-term memory coordination,
K. Zhang, F. Zhao, Y . Kang, and X. Liu, “Memory-augmented llm personalization with short-and long-term memory coordination,” arXiv preprint arXiv:2309.11696, 2023
2023 arXiv
-
[298]
Memory Matters: The Need to Improve Long-Term Memory in LLM-Agents,
K. Hatalis, D. Christou, J. Myers, S. Jones, K. Lambert, A. Amos- Binks, Z. Dannenhauer, and D. Dannenhauer, “Memory Matters: The Need to Improve Long-Term Memory in LLM-Agents,” in Proceedings of the AAAI Symposium Series , vol. 2, no. 1, 2023, pp. 277–280
2023
-
[299]
Mot: Memory-of-thought enables chatgpt to self-improve,
X. Li and X. Qiu, “Mot: Memory-of-thought enables chatgpt to self-improve,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 6354–6374
2023
-
[300]
From llm to conversational agent: A memory enhanced architecture with fine-tuning of large language models,
N. Liu, L. Chen, X. Tian, W. Zou, K. Chen, and M. Cui, “From llm to conversational agent: A memory enhanced architecture with fine-tuning of large language models,” arXiv preprint arXiv:2401.02777 , 2024. MANUSCRIPT 41
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.