Pith. sign in

REVIEW 5 major objections 6 minor 299 references

Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges

T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This survey claims that the evolution of large language models can be comprehensively mapped through three Transformer-based architectural families—encoder-only, decoder-only, and encoder-decoder—and provides a comparative analysis of…

desk verdict A useful-in-principle LLM survey, but the factual errors and duplicated copy make it unreliable in its current form. read the letter →

arxiv 2412.03220 v1 pith:ST4DGU6P submitted 2024-12-04 cs.LG

classification cs.LG
keywords LargeLanguageModels(LLMs)TransformerArchitectureGenerativeSurveyMultimodalLearningDeepNaturalProcessing(NLP)Benchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey aims to give a single, up-to-date map of large language models: how they are built, how they are trained and fine-tuned, how they are evaluated, and where they still fail. It organizes the field into three Transformer-based families—encoder-only (auto-encoding), decoder-only (auto-regressive), and encoder-decoder (sequence-to-sequence)—and traces the lineage of major models from BERT and GPT through LLaMA, PaLM, GPT-4, and multimodal systems. The intended payoff is practical: a reader could use the survey to classify any major LLM, compare it to peers on standard benchmarks, and identify the dominant techniques for adaptation and compression. If the survey is correct, it saves researchers and practitioners the effort of assembling this picture from dozens of primary sources.

What carries the argument

The organizing machinery is the three-way architectural taxonomy built on the Transformer foundation: auto-encoding (encoder-only) models such as BERT, ERNIE, and ALBERT; auto-regressive (decoder-only) models such as GPT, LLaMA, PaLM, and KOSMOS-1; and sequence-to-sequence (encoder-decoder) models such as T5, BART, Pangu, and GLM. The Transformer—a neural architecture using multi-head self-attention and positional encoding—is treated as the shared substrate, and each family is defined by which part of the Transformer it keeps and what training objective it uses. This taxonomy carries the survey's comparative analysis: benchmark tables and lineage diagrams are organized by family, and fine-tuning and compression techniques are discussed as ways of adapting models within a family.

What would settle it

Spot-check at least ten model descriptions from the survey against the cited technical reports; for example, verify the release year, parameter count, and architectural details claimed for LLaMA, Megatron, and KOSMOS-1. Any material mismatch, such as dating LLaMA to 2022 when its cited source describes the 2023 LLaMA-2 model, would indicate that the survey's secondary summaries cannot be trusted without independent verification.

Watch

Extended reading notes

Core claim

On its own terms, the paper discovers nothing new about language models; its claim is that the recent history of LLMs is now mature enough to be summarized in a coherent narrative, and that the right narrative is an architectural one. Grouping models by encoder-only, decoder-only, and encoder-decoder designs, the survey traces evolution from 2018's GPT and BERT through the open-weights LLaMA family, the Pathways-based PaLM series, and the multimodal GPT-4, KOSMOS-1, and Gemini models. It further claims that standardized benchmarks—MMLU, SuperGLUE, HellaSwag, ARC, WinoGrande for language, NLVR2 and VQA for vision-language—plus a taxonomy of fine-tuning methods (LoRA and parameter-efficient techniques) and challenges (data quality, compression, distributed training, multimodality) provide a reliable basis for comparing models and guiding practice. The contribution is the synthesis and the comparative framing, not a new model or algorithm.

Load-bearing premise

The survey's value rests on the assumption that its selection of models, benchmarks, and methods is representative and that the descriptions of cited works faithfully reflect their primary sources; if either fails, the overview can mislead despite its breadth.

Editorial extensions

If this is right

  • A reader can classify any new LLM release by its architectural family and predict its likely strengths and trade-offs, since the survey links family to task suitability.
  • Practitioners can use the benchmark figures (MMLU, HellaSwag, ARC, WinoGrande, NLVR2, VQA) to compare models released in different years without re-running evaluations.
  • The survey's catalog of LoRA and other parameter-efficient fine-tuning methods offers concrete, lower-cost routes for adapting large models to specialized tasks.
  • The explicit list of challenges—data quality and bias, model compression, distributed computation, and multimodal alignment—marks where near-term research effort is most needed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the survey's own comparison tables suggest that benchmark scores, not parameter counts, are becoming the field's common currency, but the paper does not argue for any particular evaluation protocol.
  • Editorial inference: because the paper includes a few clear factual slips (for example, dating LLaMA to 2022 and citing the LLaMA-2 paper for LLaMA-1, and duplicating the Megatron description), a careful reader should treat specific numbers and dates as leads to check against primary sources rather than as authoritative.
  • Editorial inference: the architectural taxonomy could be extended to organize future model families, but doing so would require explicit inclusion criteria, which the survey does not provide.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper is a survey of Large Language Model (LLM) and Multimodal LLM (MLLM) research, aiming to cover architectures (encoder-only, decoder-only, encoder-decoder), benchmark evaluations, fine-tuning and pre-training methods, applications, and challenges, with coverage of models up to mid-2024 and a reference list of 427 entries. It positions itself as a comprehensive, holistic map of the field, explicitly claiming in Section III to delve deeply into Architecture, Benchmark, and Challenges aspects.

Significance. If the survey were factually reliable, its breadth—spanning model evolution, PEFT taxonomies, benchmark comparisons, and challenge taxonomies—would make it a useful entry point for researchers and practitioners. The manuscript has genuine strengths: it compiles a large reference base, includes a comparative table of earlier surveys (Table 3), and offers a structured taxonomy of parameter-efficient fine-tuning methods (Table 5). However, the paper's value rests entirely on the accuracy and representativeness of its content, and the demonstrated factual errors, duplicated passages, and unverifiable benchmark figures currently undermine that value. The errors are correctable within the scope of a survey, so the appropriate path is a major revision rather than rejection, but the revision must be substantive rather than cosmetic.

major comments (5)
  1. [§VI-B-5] The section states that 'The LLaMA model was introduced by Meta in 2022' and cites reference [10], which is the LLaMA-2 paper, for claims about the original LLaMA. LLaMA-1 was released in February 2023 and has its own technical report; conflating the two releases corrupts the historical timeline that the survey claims to provide. This needs correction, including separate references for LLaMA-1 and LLaMA-2.
  2. [§VI-B-4] The Megatron paragraph is duplicated nearly verbatim: the text beginning 'Nvidia Megatron [220] is a framework proposed by Nvidia...' appears twice, with the second copy mislabeled as following from 'ChatGPT Nvidia's Megatron'. The two copies also give contradictory definitions: the first says inter-layer parallel is also known as tensor parallel, while the second says intra-layer parallelism is also known as tensor parallelism. This is a load-bearing editing failure in a technical survey, and the passage must be rewritten with a single, correct definition of tensor, pipeline, and data parallelism.
  3. [§VI-B-1, §VI-C-1, §VI-B-7] Multiple model descriptions contain factual errors that are not isolated typos. In §VI-B-1, GPT is described as having 'employed a 12-layer transformer encoder,' but GPT is decoder-only, contradicting the paper's own classification in §II-C. In §VI-C-1, BART is attributed to 'the Google research team,' whereas BART is from Facebook AI Research. In §VI-B-7, CogView is described as 'a 4T parameter Chinese multimodal LLM' (CogView is a 4B-parameter text-to-image model), and Mistral 7B is said to 'use mix-of-expert,' which is actually a property of Mixtral, not Mistral 7B. Together with the LLaMA error, these indicate a systemic reliability problem in the model descriptions, which is the core content of a survey.
  4. [§IV-A, Figures 5–11] The benchmark performance figures plot named models (MMLU, HellaSwag, ARC, WinoGrande, NLVR2, VQA) without any accompanying data tables, source citations, or evaluation-protocol details. The reader cannot verify the plotted values, determine whether they come from the original benchmark papers or secondary leaderboards, or assess comparability across models with different prompting and few-shot settings. Since the paper explicitly claims in §III to delve deeply into benchmarking, these unsupported figures are load-bearing for that claim and must be replaced or supplemented with tables that give exact scores, sources, and protocols.
  5. [§III, §IV] The paper does not state inclusion criteria for the models, benchmarks, or prior surveys it discusses. The claim of being 'comprehensive' and 'holistic' (abstract and §III) is therefore unsupported by any reproducible methodology. To make the survey's coverage verifiable, the authors should specify how models and benchmark results were selected (e.g., date ranges, release venues, availability of primary sources) and how the comparative table of prior surveys (Table 3) was compiled.
minor comments (6)
  1. [§VI-B-5, References] Reference [10] is cited for both LLaMA and LLaMA-2 claims; these should be separate references, with the original LLaMA report cited for LLaMA-1.
  2. [§II-A, Equations (1)–(2)] The positional-encoding formulas are printed as '100002i/dim' in the denominator; this should be typeset as 10000^(2i/dim) to be unambiguous.
  3. [§VIII-A-3, §VIII-B-1] Cross-references are incorrect: the text says 'figure 15 shows the common source of the datasets' but the relevant figure is Figure 28, and 'Table 7 shows transformer based pruning technology' but the pruning methods are in Table 9.
  4. [§VI-B-4] The phrase 'ChatGPT Nvidia's Megatron' appears in the running text as a leftover editing artifact and should be removed.
  5. [Throughout] There are numerous typos and terminological inconsistencies, including 'Megatron-Tuiring NLG' (§VI-B-4), 'the the datasets' (§VIII-A-3), and inconsistent spelling of 'auto-regressive' vs. 'autoregressive' and 'LLaMA' vs. 'Llama'.
  6. [Figures 12–14] The evolutionary tree and parameter-count figures lack explicit sources and some lack axis labels (e.g., Figure 13 uses a log scale without stating the base); these should be clarified for a reader to interpret the data.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a survey with no derived predictions, fitted parameters, or load-bearing self-citation chain.

full rationale

This manuscript is a literature survey of LLM and MLLM architectures, benchmarks, fine-tuning techniques, challenges, and applications. It does not propose a new model, derive a mathematical result, fit parameters to data, or make a prediction that is then validated against the same data. The central claims are descriptive: that the survey organizes prior work and that its coverage is comprehensive. Those claims rest on the accuracy and representativeness of the cited sources, not on any equation that defines a target quantity in terms of the paper's own outputs. There is no self-definitional step, no fitted input renamed as a prediction, and no invoked uniqueness theorem or ansatz imported from prior work by the same authors. The paper's known weaknesses—such as the LLaMA 1/LLaMA 2 dating and citation conflation in Section VI-B-5, the duplicated Megatron paragraph in Section VI-B-4, and unsourced benchmark figures—are factual reliability and correctness risks, not circular reasoning. A survey can be inaccurate without being circular. Accordingly, no circularity steps are identified, and the score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This review introduces no free parameters and no invented entities. Its central claim rests on the adequacy of its taxonomy, the representativeness of its benchmark selection, the correctness of its descriptions of cited works, and the accuracy of its figures. These are domain assumptions, not mathematically derivable facts.

assumptions (3)
  • domain assumption The taxonomy of LLMs into encoder-only, decoder-only, and encoder-decoder architectures is complete enough to organize all relevant modern models.
    Section II and Section VI classify the entire model landscape with this trichotomy; if a major model family does not fit, the survey's organizing principle weakens.
  • domain assumption The benchmarks selected in Section IV (MMLU, SuperGLUE, HellaSwag, ARC, WinoGrande, NLVR2, VQA) are representative of how LLM capability should be measured.
    No inclusion criteria are given for choosing these benchmarks over others, yet the comparative analysis is built on them.
  • domain assumption Figures 5 through 14 display correct scores sourced from credible leaderboards or papers, meaning the displayed data accurately represent the cited models.
    The figure captions do not state data sources or extraction dates, so the charts are unverifiable from the text alone.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges." pith.science (2026). https://pith.science/paper/ST4DGU6P

@misc{pith2026241203220,
  author       = {Pith},
  title        = {Pith review of: Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ST4DGU6P}},
  note         = {Machine review of arXiv:2412.03220}
}
read the original abstract

Large Language Models (LLMs) represent a class of deep learning models adept at understanding natural language and generating coherent responses to various prompts or queries. These models far exceed the complexity of conventional neural networks, often encompassing dozens of neural network layers and containing billions to trillions of parameters. They are typically trained on vast datasets, utilizing architectures based on transformer blocks. Present-day LLMs are multi-functional, capable of performing a range of tasks from text generation and language translation to question answering, as well as code generation and analysis. An advanced subset of these models, known as Multimodal Large Language Models (MLLMs), extends LLM capabilities to process and interpret multiple data modalities, including images, audio, and video. This enhancement empowers MLLMs with capabilities like video editing, image comprehension, and captioning for visual content. This survey provides a comprehensive overview of the recent advancements in LLMs. We begin by tracing the evolution of LLMs and subsequently delve into the advent and nuances of MLLMs. We analyze emerging state-of-the-art MLLMs, exploring their technical features, strengths, and limitations. Additionally, we present a comparative analysis of these models and discuss their challenges, potential limitations, and prospects for future development.

Figures

Figures reproduced from arXiv: 2412.03220 by the authors.

Figure 1
Figure 1. FIGURE 1 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIGURE 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. shows the difference between of the principle of auto-encoders and variational auto-encoders. Color, Shape, Size, Height, ... Feature Attributes Input Reconstruction Color Shape Size Height A Encoder Decoder Feature Distribution Input Sample & Generate B Encoder Decoder ... FIGURE 3. (A) Workflow of Auto-encoder, auto-encoder encode the feature attribute directly. (B) Workflow of Variational Auto-encoder, different … view at source ↗
Figures from the paper (20 more)
Figure 4
Figure 4. Figure 4: FIGURE 4 [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: FIGURE 5 [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: FIGURE 6 [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 9
Figure 9. Figure 9: FIGURE 9 [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 8
Figure 8. Figure 8: FIGURE 8 [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 11
Figure 11. Figure 11: FIGURE 11 [PITH_FULL_IMAGE:figures/full_fig_p009_11.png]
Figure 12
Figure 12. Figure 12: FIGURE 12 [PITH_FULL_IMAGE:figures/full_fig_p011_12.png]
Figure 13
Figure 13. Figure 13: FIGURE 13 [PITH_FULL_IMAGE:figures/full_fig_p012_13.png]
Figure 14
Figure 14. Figure 14: FIGURE 14 [PITH_FULL_IMAGE:figures/full_fig_p012_14.png]
Figure 17
Figure 17. Figure 17: FIGURE 17 [PITH_FULL_IMAGE:figures/full_fig_p013_17.png]
Figure 16
Figure 16. Figure 16: FIGURE 16 [PITH_FULL_IMAGE:figures/full_fig_p013_16.png]
Figure 19
Figure 19. Figure 19: FIGURE 19 [PITH_FULL_IMAGE:figures/full_fig_p014_19.png]
Figure 18
Figure 18. Figure 18: FIGURE 18 [PITH_FULL_IMAGE:figures/full_fig_p014_18.png]
Figure 20
Figure 20. Figure 20: FIGURE 20 [PITH_FULL_IMAGE:figures/full_fig_p015_20.png]
Figure 23
Figure 23. Figure 23: FIGURE 23 [PITH_FULL_IMAGE:figures/full_fig_p017_23.png]
Figure 24
Figure 24. Figure 24: FIGURE 24 [PITH_FULL_IMAGE:figures/full_fig_p018_24.png]
Figure 26
Figure 26. Figure 26: FIGURE 26 [PITH_FULL_IMAGE:figures/full_fig_p022_26.png]
Figure 27
Figure 27. Figure 27: FIGURE 27 [PITH_FULL_IMAGE:figures/full_fig_p023_27.png]
Figure 28
Figure 28. Figure 28: FIGURE 28 [PITH_FULL_IMAGE:figures/full_fig_p025_28.png]
Figure 30
Figure 30. Figure 30: FIGURE 30 [PITH_FULL_IMAGE:figures/full_fig_p029_30.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

299 extracted references · 4 canonical work pages

  1. [10]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P . Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P . Bhargava, S. Bhosale et al. , ‘‘Llama 2: Open foundation and fine-tuned chat models,’’ arXiv preprint arXiv:2307.09288, 2023

  2. [220]

    [Online]

    NVIDIA, ‘‘Nvidia nemo_2023,’’ Oct 2023. [Online]. Available: https://docs.nvidia.com/deeplearning/nemo/user-guide/docs/en/main/ nlp/megatron.html

  3. [1]

    V aswani, N

    A. V aswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, ‘‘Attention is all you need,’’ Advances in neural information processing systems, vol. 30, 2017

  4. [2]

    Radford, K

    A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al., ‘‘Improving language understanding by generative pre-training,’’ 2018

  5. [3]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, ‘‘Bert: Pre-training of deep bidirectional transformers for language understanding,’’ arXiv preprint arXiv:1810.04805, 2018

  6. [4]

    W. Zeng, X. Ren, T. Su, H. Wang, Y . Liao, Z. Wang, X. Jiang, Z. Y ang, K. Wang, X. Zhang et al., ‘‘Pangu-{\alpha}: Large-scale autoregressive pretrained chinese language models with auto-parallel computation,’’ arXiv preprint arXiv:2104.12369, 2021

  7. [5]

    X. Ren, P . Zhou, X. Meng, X. Huang, Y . Wang, W. Wang, P . Li, X. Zhang, A. Podolskiy, G. Arshinov et al. , ‘‘Pangu-{\Sigma}: Towards trillion parameter language model with sparse heterogeneous computing,’’ arXiv preprint arXiv:2303.10845, 2023

  8. [6]

    Graves and A

    A. Graves and A. Graves, ‘‘Long short-term memory,’’ Supervised sequence labelling with recurrent neural networks , pp. 37–45, 2012

Show all 299 references
  1. [7]

    L. R. Medsker and L. Jain, ‘‘Recurrent neural networks,’’ Design and Applications, vol. 5, no. 64-67, p. 2, 2001

  2. [8]

    Zhang, X

    Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu, ‘‘Ernie: Enhanced language representation with informative entities,’’ arXiv preprint arXiv:1905.07129, 2019

  3. [9]

    Z. Lan, M. Chen, S. Goodman, K. Gimpel, P . Sharma, and R. Soricut, ‘‘Albert: A lite bert for self-supervised learning of language representations,’’ arXiv preprint arXiv:1909.11942, 2019

  4. [11]

    Raffel, N

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P . J. Liu, ‘‘Exploring the limits of transfer learning with a unified text-to-text transformer,’’ The Journal of Machine Learning Research, vol. 21, no. 1, pp. 5485–5551, 2020

  5. [12]

    Z. Du, Y . Qian, X. Liu, M. Ding, J. Qiu, Z. Y ang, and J. Tang, ‘‘Glm: General language model pretraining with autoregressive blank infilling,’’ arXiv preprint arXiv:2103.10360, 2021. 30 VOLUME 11, 2024 Minghao et al.: Survey of different Large Language Model Architectures: T...

  6. [13]

    D. P . Kingma, ‘‘Auto-encoding variational bayes,’’ arXiv preprint arXiv:1312.6114, 2013

  7. [14]

    J. Zhai, S. Zhang, J. Chen, and Q. He, ‘‘Autoencoder and its various variants,’’ in 2018 IEEE international conference on systems, man, and cybernetics (SMC). IEEE, 2018, pp. 415–419

  8. [15]

    R. Wei, C. Garcia, A. El-Sayed, V . Peterson, and A. Mahmood, ‘‘V ariations in variational autoencoders-a comparative evaluation,’’ Ieee Access, vol. 8, pp. 153 651–153 670, 2020

  9. [16]

    I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, ‘‘Generative adversarial networks,’’ Science Robotics, vol. 3, pp. 2672–2680, 6 2014

  10. [17]

    Radford, L

    A. Radford, L. Metz, and S. Chintala, ‘‘Unsupervised representation learning with deep convolutional generative adversarial networks,’’ 4th International Conference on Learning Representations, ICLR 2016 - Conference Track Proceedings, 11 2015

  11. [18]

    Mirza and S

    M. Mirza and S. Osindero, ‘‘Conditional generative adversarial nets,’’ arXiv preprint: arXiv:1411.1784, 11 2014

  12. [19]

    Arjovsky, S

    M. Arjovsky, S. Chintala, and L. Bottou, ‘‘Wasserstein gan,’’ arXiv preprint: arXiv:1701.07875, 1 2017

  13. [20]

    Gulrajani, F

    I. Gulrajani, F. Ahmed, M. Arjovsky, V . Dumoulin, and A. Courville, ‘‘Improved training of wasserstein gans,’’ Advances in Neural Informa- tion Processing Systems, vol. 2017-December, pp. 5768–5778, 3 2017

  14. [21]

    X. Mao, Q. Li, H. Xie, R. Y . Lau, Z. Wang, and S. P . Smolley, ‘‘Least squares generative adversarial networks,’’ Proceedings of the IEEE International Conference on Computer Vision , vol. 2017-October, pp. 2813–2821, 11 2016

  15. [22]

    J. Y . Zhu, T. Park, P . Isola, and A. A. Efros, ‘‘Unpaired image-to-image translation using cycle-consistent adversarial networks,’’ Proceedings of the IEEE International Conference on Computer Vision , vol. 2017-October, pp. 2242–2251, 3 2017

  16. [23]

    Karras, S

    T. Karras, S. Laine, and T. Aila, ‘‘A style-based generator architecture for generative adversarial networks,’’IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, pp. 4217–4228, 12 2018

  17. [24]

    Zhang, I

    H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena, ‘‘Self-attention gen- erative adversarial networks,’’36th International Conference on Machine Learning, ICML 2019, vol. 2019-June, pp. 12 744–12 753, 5 2018

  18. [25]

    Brock, J

    A. Brock, J. Donahue, and K. Simonyan, ‘‘Large scale gan training for high fidelity natural image synthesis,’’ 7th International Conference on Learning Representations, ICLR 2019 , 9 2018

  19. [26]

    Karras, T

    T. Karras, T. Aila, S. Laine, and J. Lehtinen, ‘‘Progressive growing of gans for improved quality, stability, and variation,’’ 6th International Conference on Learning Representations, ICLR 2018 - Conference Track Proceedings, 10 2017

  20. [27]

    Y . Choi, M. Choi, M. Kim, J. W. Ha, S. Kim, and J. Choo, ‘‘Stargan: Uni- fied generative adversarial networks for multi-domain image-to-image translation,’’ Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 8789–8797, 11 2017

  21. [28]

    Z. Wang, Q. She, and T. E. Ward, ‘‘Generative adversarial networks in computer vision: A survey and taxonomy,’’ ACM Computing Surveys , vol. 54, 6 2019

  22. [29]

    J. Gui, Z. Sun, Y . Wen, D. Tao, and J. Y e, ‘‘A review on generative adversarial networks: Algorithms, theory, and applications,’’ IEEE Transactions on Knowledge and Data Engineering , vol. 35, pp. 3313–3332, 4 2023

  23. [30]

    Creswell, T

    A. Creswell, T. White, V . Dumoulin, K. Arulkumaran, B. Sengupta, and A. A. Bharath, ‘‘Generative adversarial networks: An overview,’’ IEEE Signal Processing Magazine, vol. 35, pp. 53–65, 10 2017

  24. [31]

    I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Der- noncourt, T. Y u, R. Zhang, and N. K. Ahmed, ‘‘Bias and fairness in large language models: A survey,’’ arXiv preprint arXiv:2309.00770, 2023

  25. [32]

    H. Shi, Z. Xu, H. Wang, W. Qin, W. Wang, Y . Wang, and H. Wang, ‘‘Continual learning of large language models: A comprehensive survey,’’ arXiv preprint arXiv:2404.16789, 2024

  26. [33]

    Z. Zhou, X. Ning, K. Hong, T. Fu, J. Xu, S. Li, Y . Lou, L. Wang, Z. Y uan, X. Li et al., ‘‘A survey on efficient inference for large language models,’’ arXiv preprint arXiv:2404.14294, 2024

  27. [34]

    X. Fang, W. Xu, F. A. Tan, J. Zhang, Z. Hu, Y . Qi, S. Nickleach, D. Socolinsky, S. Sengamedu, and C. Faloutsos, ‘‘Large language models (llms) on tabular data: Prediction, generation, and understanding–a survey,’’arXiv preprint arXiv:2402.17944, 2024

  28. [35]

    B. C. Das, M. H. Amini, and Y . Wu, ‘‘Security and privacy challenges of large language models: A survey,’’ arXiv preprint arXiv:2402.00888 , 2024

  29. [36]

    H. Jin, L. Hu, X. Li, P . Zhang, C. Chen, J. Zhuang, and H. Wang, ‘‘Jail- breakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models,’’ arXiv preprint arXiv:2407.01599, 2024

  30. [37]

    L. Wu, Z. Zheng, Z. Qiu, H. Wang, H. Gu, T. Shen, C. Qin, C. Zhu, H. Zhu, Q. Liu et al., ‘‘A survey on large language models for recommendation,’’ arXiv preprint arXiv:2305.19860, 2023

  31. [38]

    Akyash and H

    M. Akyash and H. M. Kamali, ‘‘Evolutionary large language models for hardware security: A comparative survey,’’ arXiv preprint arXiv:2404.16651, 2024

  32. [39]

    T. Bai, H. Liang, B. Wan, L. Y ang, B. Li, Y . Wang, B. Cui, C. He, B. Y uan, and W. Zhang, ‘‘A survey of multimodal large language model from a data-centric perspective,’’ arXiv preprint arXiv:2405.16640, 2024

  33. [40]

    H. Xiao, F. Zhou, X. Liu, T. Liu, Z. Li, X. Liu, and X. Huang, ‘‘A comprehensive survey of large language models and multimodal large language models in medicine,’’ arXiv preprint arXiv:2405.08603, 2024

  34. [41]

    H. Zhou, C. Hu, Y . Y uan, Y . Cui, Y . Jin, C. Chen, H. Wu, D. Y uan, L. Jiang, D. Wu et al. , ‘‘Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities,’’ arXiv preprint arXiv:2405.10825, 2024

  35. [42]

    Huang, K

    Y . Huang, K. Tang, and M. Chen, ‘‘A comprehensive survey on evaluating large language model applications in the medical industry,’’ arXiv preprint arXiv:2404.15777, 2024

  36. [43]

    C. Qu, S. Dai, X. Wei, H. Cai, S. Wang, D. Yin, J. Xu, and J.-R. Wen, ‘‘Tool learning with large language models: A survey,’’ arXiv preprint arXiv:2405.17935, 2024

  37. [44]

    L. Qin, Q. Chen, X. Feng, Y . Wu, Y . Zhang, Y . Li, M. Li, W. Che, and P . S. Y u, ‘‘Large language models meet nlp: A survey,’’ arXiv preprint arXiv:2405.12819, 2024

  38. [45]

    Zhang, Y

    Z. Zhang, Y . Sun, Z. Wang, Y . Nie, X. Ma, P . Sun, and R. Li, ‘‘Large language models for mobility in transportation systems: A survey on forecasting tasks,’’ arXiv preprint arXiv:2405.02357, 2024

  39. [46]

    Kukreja, T

    S. Kukreja, T. Kumar, A. Purohit, A. Dasgupta, and D. Guha, ‘‘A literature survey on open source large language models,’’ in Proceedings of the 2024 7th International Conference on Computers in Management and Business, 2024, pp. 133–143

  40. [47]

    L. Qin, Q. Chen, Y . Zhou, Z. Chen, Y . Li, L. Liao, M. Li, W. Che, and P . S. Y u, ‘‘Multilingual large language model: A survey of resources, taxonomy and frontiers,’’ arXiv preprint arXiv:2404.04925, 2024

  41. [48]

    S. Dai, C. Xu, S. Xu, L. Pang, Z. Dong, and J. Xu, ‘‘Unifying bias and un- fairness in information retrieval: A survey of challenges and opportunities with large language models,’’ arXiv preprint arXiv:2404.11457, 2024

  42. [49]

    S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, ‘‘A survey on multimodal large language models,’’ arXiv preprint arXiv:2306.13549 , 2023

  43. [50]

    Zhang, X

    Z. Zhang, X. Bo, C. Ma, R. Li, X. Chen, Q. Dai, J. Zhu, Z. Dong, and J.-R. Wen, ‘‘A survey on the memory mechanism of large language model based agents,’’ arXiv preprint arXiv:2404.13501, 2024

  44. [51]

    S. Hu, T. Huang, F. Ilhan, S. Tekin, G. Liu, R. Kompella, and L. Liu, ‘‘A survey on large language model-based game agents,’’ arXiv preprint arXiv:2404.02039, 2024

  45. [52]

    Z. Bai, P . Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou, ‘‘Hallucination of multimodal large language models: A survey,’’ arXiv preprint arXiv:2404.18930, 2024

  46. [53]

    Huang and J

    Y . Huang and J. Huang, ‘‘A survey on retrieval-augmented text generation for large language models,’’ arXiv preprint arXiv:2404.10981, 2024

  47. [54]

    J. Li, T. Tang, W. X. Zhao, J.-Y . Nie, and J.-R. Wen, ‘‘Pre-trained language models for text generation: A survey,’’ ACM Computing Surveys, vol. 56, no. 9, pp. 1–39, 2024

  48. [55]

    Y . Liu, Y . Y ao, J.-F. Ton, X. Zhang, R. G. H. Cheng, Y . Klochkov, M. F. Taufiq, and H. Li, ‘‘Trustworthy llms: a survey and guideline for evaluating large language models’ alignment,’’ arXiv preprint arXiv:2308.05374, 2023

  49. [56]

    Y . Y ao, J. Duan, K. Xu, Y . Cai, Z. Sun, and Y . Zhang, ‘‘A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,’’High-Confidence Computing, p. 100211, 2024

  50. [57]

    Chang, X

    Y . Chang, X. Wang, J. Wang, Y . Wu, L. Y ang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang et al. , ‘‘A survey on evaluation of large language models,’’ ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 3, pp. 1–45, 2024

  51. [58]

    Y . Gao, Y . Xiong, X. Gao, K. Jia, J. Pan, Y . Bi, Y . Dai, J. Sun, and H. Wang, ‘‘Retrieval-augmented generation for large language models: A survey,’’arXiv preprint arXiv:2312.10997, 2023. VOLUME 11, 2024 31 Minghao et al.: Survey of different Large Language Model Architect...

  52. [59]

    Zhang, L

    S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu et al., ‘‘Instruction tuning for large language models: A survey,’’arXiv preprint arXiv:2308.10792, 2023

  53. [60]

    R. Hong, X. Pang, and C. Zhang, ‘‘Advances in reasoning by prompting large language models: A survey,’’ Cybernetics and Intelligence, 2024

  54. [61]

    B. Y an, K. Li, M. Xu, Y . Dong, Y . Zhang, Z. Ren, and X. Cheng, ‘‘On protecting the data privacy of large language models (llms): A survey,’’ arXiv preprint arXiv:2403.05156, 2024

  55. [62]

    Y . Cao, H. Zhao, Y . Cheng, T. Shu, G. Liu, G. Liang, J. Zhao, and Y . Li, ‘‘Survey on large language model-enhanced reinforcement learning: Con- cept, taxonomy, and methods,’’ arXiv preprint arXiv:2404.00282, 2024

  56. [63]

    X. Liu, P . Xu, J. Wu, J. Y uan, Y . Y ang, Y . Zhou, F. Liu, T. Guan, H. Wang, T. Y uet al., ‘‘Large language models and causal inference in collabora- tion: A comprehensive survey,’’ arXiv preprint arXiv:2403.09606, 2024

  57. [64]

    Esmradi, D

    A. Esmradi, D. W. Yip, and C. F. Chan, ‘‘A comprehensive survey of attack techniques, implementation, and mitigation strategies in large language models,’’ in International Conference on Ubiquitous Security . Springer, 2023, pp. 76–95

  58. [65]

    A. G. Chowdhury, M. M. Islam, V . Kumar, F. H. Shezan, V . Jain, and A. Chadha, ‘‘Breaking down the defenses: A comparative survey of at- tacks on large language models,’’arXiv preprint arXiv:2403.04786, 2024

  59. [66]

    Sun, ‘‘A short survey of viewing large language models in legal aspect,’’ arXiv preprint arXiv:2303.09136, 2023

    Z. Sun, ‘‘A short survey of viewing large language models in legal aspect,’’ arXiv preprint arXiv:2303.09136, 2023

  60. [67]

    H. Zhao, H. Chen, F. Y ang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, and M. Du, ‘‘Explainability for large language models: A survey,’’ ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 2, pp. 1–38, 2024

  61. [68]

    Y . Zhu, H. Y uan, S. Wang, J. Liu, W. Liu, C. Deng, Z. Dou, and J.-R. Wen, ‘‘Large language models for information retrieval: A survey,’’arXiv preprint arXiv:2308.07107, 2023

  62. [69]

    J. Li, Y . Liu, C. Liu, L. Shi, X. Ren, Y . Zheng, Y . Liu, and Y . Xue, ‘‘A cross-language investigation into jailbreak attacks in large language models,’’ arXiv preprint arXiv:2401.16765, 2024

  63. [70]

    Z. Xi, W. Chen, X. Guo, W. He, Y . Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou et al. , ‘‘The rise and potential of large language model based agents: A survey,’’ arXiv preprint arXiv:2309.07864, 2023

  64. [71]

    Huang, W

    L. Huang, W. Y u, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin et al. , ‘‘A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,’’ arXiv preprint arXiv:2311.05232, 2023

  65. [72]

    Shayegani, M

    E. Shayegani, M. A. A. Mamun, Y . Fu, P . Zaree, Y . Dong, and N. Abu-Ghazaleh, ‘‘Survey of vulnerabilities in large language models revealed by adversarial attacks,’’arXiv preprint arXiv:2310.10844, 2023

  66. [73]

    Zhang, H

    H. Zhang, H. Song, S. Li, M. Zhou, and D. Song, ‘‘A survey of controllable text generation using transformer-based pre-trained language models,’’ ACM Computing Surveys, vol. 56, no. 3, pp. 1–37, 2023

  67. [74]

    B. Min, H. Ross, E. Sulem, A. P . B. V eyseh, T. H. Nguyen, O. Sainz, E. Agirre, I. Heintz, and D. Roth, ‘‘Recent advances in natural language processing via large pre-trained language models: A survey,’’ ACM Computing Surveys, vol. 56, no. 2, pp. 1–40, 2023

  68. [75]

    Zhang, Y

    Y . Zhang, Y . Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y . Zhang, Y . Chenet al., ‘‘Siren’s song in the ai ocean: a survey on hallucination in large language models,’’ arXiv preprint arXiv:2309.01219, 2023

  69. [76]

    X. Zhu, J. Li, Y . Liu, C. Ma, and W. Wang, ‘‘A survey on model compres- sion for large language models,’’ arXiv preprint arXiv:2308.07633, 2023

  70. [77]

    L. Hu, Z. Liu, Z. Zhao, L. Hou, L. Nie, and J. Li, ‘‘A survey of knowledge enhanced pre-trained language models,’’ IEEE Transactions on Knowledge and Data Engineering, vol. 36, no. 4, pp. 1413–1430, 2024

  71. [78]

    B. Wang, Q. Xie, J. Pei, Z. Chen, P . Tiwari, Z. Li, and J. Fu, ‘‘Pre-trained language models in biomedical domain: A systematic survey,’’ ACM Computing Surveys, vol. 56, no. 3, pp. 1–52, 2023

  72. [79]

    Huang and K

    J. Huang and K. C.-C. Chang, ‘‘Towards reasoning in large language models: A survey,’’ arXiv preprint arXiv:2212.10403, 2022

  73. [80]

    Kasneci, K

    E. Kasneci, K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. Günnemann, E. Hüllermeier et al. , ‘‘Chatgpt for good? on opportunities and challenges of large language models for education,’’ Learning and individual differences , vol. 103, p...

  74. [81]

    W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong et al. , ‘‘A survey of large language models,’’ arXiv preprint arXiv:2303.18223, 2023

  75. [82]

    Mialon, R

    G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Y u, A. Celikyilmazet al., ‘‘Augmented language models: a survey,’’ arXiv preprint arXiv:2302.07842, 2023

  76. [83]

    K. S. Kalyan, A. Rajasekharan, and S. Sangeetha, ‘‘Ammu: a survey of transformer-based biomedical pretrained language models,’’ Journal of biomedical informatics, vol. 126, p. 103982, 2022

  77. [84]

    ——, ‘‘Ammus: A survey of transformer-based pretrained models in natural language processing,’’ arXiv preprint arXiv:2108.05542, 2021

  78. [85]

    M. Zaib, Q. Z. Sheng, and W. Emma Zhang, ‘‘A short survey of pre-trained language models for conversational ai-a new age in nlp,’’ in Proceedings of the Australasian computer science week multiconference , 2020, pp. 1–4

  79. [86]

    Houlsby, A

    N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, ‘‘Parameter-efficient transfer learning for nlp,’’ in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799

  80. [87]

    J. He, C. Zhou, X. Ma, T. Berg-Kirkpatrick, and G. Neubig, ‘‘Towards a unified view of parameter-efficient transfer learning,’’ arXiv preprint arXiv:2110.04366, 2021

  81. [88]

    Y . Zhu, J. Feng, C. Zhao, M. Wang, and L. Li, ‘‘Counter-interference adapter for multilingual machine translation,’’ arXiv preprint arXiv:2104.08154, 2021

  82. [89]

    T. Lei, J. Bai, S. Brahma, J. Ainslie, K. Lee, Y . Zhou, N. Du, V . Y . Zhao, Y . Wu, B. Li et al. , ‘‘Conditional adapters: Parameter-efficient transfer learning with fast inference,’’ arXiv preprint arXiv:2304.04947, 2023

  83. [90]

    Pfeiffer, A

    J. Pfeiffer, A. Kamath, A. Rücklé, K. Cho, and I. Gurevych, ‘‘Adapterfusion: Non-destructive task composition for transfer learning,’’ arXiv preprint arXiv:2005.00247, 2020

  84. [91]

    Y . Wang, S. Mukherjee, X. Liu, J. Gao, A. H. Awadallah, and J. Gao, ‘‘Adamix: Mixture-of-adapter for parameter-efficient tuning of large lan- guage models,’’arXiv preprint arXiv:2205.12410, vol. 1, no. 2, p. 4, 2022

  85. [92]

    H. Zhao, J. Fu, and Z. He, ‘‘Prototype-based hyperadapter for sample- efficient multi-task tuning,’’ arXiv preprint arXiv:2310.11670, 2023

  86. [93]

    Chronopoulou, M

    A. Chronopoulou, M. E. Peters, A. Fraser, and J. Dodge, ‘‘Adaptersoup: Weight averaging to improve generalization of pretrained language models,’’ arXiv preprint arXiv:2302.07027, 2023

  87. [94]

    He, R.-Z

    S. He, R.-Z. Fan, L. Ding, L. Shen, T. Zhou, and D. Tao, ‘‘Mera: Merging pretrained adapters for few-shot learning,’’ arXiv preprint arXiv:2308.15982, 2023

  88. [95]

    R. K. Mahabadi, S. Ruder, M. Dehghani, and J. Henderson, ‘‘Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks,’’arXiv preprint arXiv:2106.04489, 2021

  89. [96]

    X. L. Li and P . Liang, ‘‘Prefix-tuning: Optimizing continuous prompts for generation,’’ arXiv preprint arXiv:2101.00190, 2021

  90. [97]

    J. Li, W. Aitken, R. Bhambhoria, and X. Zhu, ‘‘Prefix propagation: Parameter-efficient tuning for long sequences,’’ arXiv preprint arXiv:2305.12086, 2023

  91. [98]

    X. Liu, K. Ji, Y . Fu, W. L. Tam, Z. Du, Z. Y ang, and J. Tang, ‘‘P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks,’’ arXiv preprint arXiv:2110.07602, 2021

  92. [99]

    Zhang, C

    Z.-R. Zhang, C. Tan, H. Xu, C. Wang, J. Huang, and S. Huang, ‘‘Towards adaptive prefix tuning for parameter-efficient language model fine-tuning,’’ arXiv preprint arXiv:2305.15212, 2023

  93. [100]

    X. Liu, Y . Zheng, Z. Du, M. Ding, Y . Qian, Z. Y ang, and J. Tang, ‘‘Gpt understands, too,’’ arXiv preprint arXiv:2103.10385, 2021

  94. [101]

    Lester, R

    B. Lester, R. Al-Rfou, and N. Constant, ‘‘The power of scale for parameter-efficient prompt tuning,’’ arXiv preprint arXiv:2104.08691 , 2021

  95. [102]

    F. Ma, C. Zhang, L. Ren, J. Wang, Q. Wang, W. Wu, X. Quan, and D. Song, ‘‘Xprompt: Exploring the extreme of prompt tuning,’’ arXiv preprint arXiv:2210.04457, 2022

  96. [103]

    Z. Wu, S. Wang, J. Gu, R. Hou, Y . Dong, V . Vydiswaran, and H. Ma, ‘‘Idpg: An instance-dependent prompt generation method,’’ arXiv preprint arXiv:2204.04497, 2022

  97. [104]

    X. Liu, T. Sun, X. Huang, and X. Qiu, ‘‘Late prompt tuning: A late prompt could be better than many prompts,’’ arXiv preprint arXiv:2210.11292 , 2022

  98. [105]

    Zhu and M

    W. Zhu and M. Tan, ‘‘Spt: Learning to selectively insert prompts for better prompt tuning,’’ in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 11 862–11 878

  99. [106]

    Q. Wang, Y . Mao, J. Wang, H. Y u, S. Nie, S. Wang, F. Feng, L. Huang, X. Quan, Z. Xu et al. , ‘‘Aprompt: Attention prompt tuning for efficient adaptation of pre-trained language models,’’ in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processin...

  100. [107]

    T. Vu, B. Lester, N. Constant, R. Al-Rfou, and D. Cer, ‘‘Spot: Better frozen model adaptation through soft prompt transfer,’’ arXiv preprint arXiv:2110.07904, 2021

  101. [108]

    Y . Su, X. Wang, Y . Qin, C.-M. Chan, Y . Lin, H. Wang, K. Wen, Z. Liu, P . Li, J. Li et al. , ‘‘On transferability of prompt tuning for natural language processing,’’ arXiv preprint arXiv:2111.06719, 2021

  102. [109]

    J. Wu, T. Y u, R. Wang, Z. Song, R. Zhang, H. Zhao, C. Lu, S. Li, and R. Henao, ‘‘Infoprompt: Information-theoretic soft prompt tuning for nat- ural language understanding,’’ arXiv preprint arXiv:2306.04933, 2023

  103. [110]

    L. Chen, H. Huang, and M. Cheng, ‘‘Ptp: Boosting stability and performance of prompt tuning with perturbation-based regularizer,’’ arXiv preprint arXiv:2305.02423, 2023

  104. [111]

    Y . Qin, X. Wang, Y . Su, Y . Lin, N. Ding, J. Yi, W. Chen, Z. Liu, J. Li, L. Hou et al. , ‘‘Exploring universal intrinsic task subspace via prompt tuning,’’ arXiv preprint arXiv:2110.07867, 2021

  105. [112]

    J.-Y . Choi, J. Kim, J.-H. Park, W.-L. Mok, and S. Lee, ‘‘Smop: Towards efficient and effective prompt tuning with sparse mixture-of-prompts,’’ in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 14 306–14 316

  106. [113]

    Shi and A

    Z. Shi and A. Lipani, ‘‘Dept: Decomposed prompt tuning for parameter-efficient fine-tuning,’’arXiv preprint arXiv:2309.05173, 2023

  107. [114]

    H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, and C. A. Raffel, ‘‘Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 1950–1965, 2022

  108. [115]

    Zadouri, A

    T. Zadouri, A. Üstün, A. Ahmadian, B. Ermiş, A. Locatelli, and S. Hooker, ‘‘Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning,’’ arXiv preprint arXiv:2309.05444, 2023

  109. [116]

    D. Lian, D. Zhou, J. Feng, and X. Wang, ‘‘Scaling & shifting your features: A new baseline for efficient model tuning,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 109–123, 2022

  110. [117]

    X. Lu, F. Brahman, P . West, J. Jang, K. Chandu, A. Ravichander, L. Qin, P . Ammanabrolu, L. Jiang, S. Ramnath et al. , ‘‘Inference-time policy adapters (ipa): Tailoring extreme-scale lms without fine-tuning,’’ arXiv preprint arXiv:2305.15065, 2023

  111. [118]

    D. Guo, A. M. Rush, and Y . Kim, ‘‘Parameter-efficient transfer learning with diff pruning,’’ arXiv preprint arXiv:2012.07463, 2020

  112. [119]

    Lawton, A

    N. Lawton, A. Kumar, G. Thattai, A. Galstyan, and G. V . Steeg, ‘‘Neural architecture search for parameter-efficient fine-tuning of large pre-trained language models,’’ arXiv preprint arXiv:2305.16597, 2023

  113. [120]

    B. Liao, Y . Meng, and C. Monz, ‘‘Parameter-efficient fine-tuning without introducing new latency,’’ arXiv preprint arXiv:2305.16742, 2023

  114. [121]

    Y .-L. Sung, V . Nair, and C. A. Raffel, ‘‘Training neural networks with fixed sparse masks,’’ Advances in Neural Information Processing Systems, vol. 34, pp. 24 193–24 205, 2021

  115. [122]

    S. S. S. Das, R. H. Zhang, P . Shi, W. Yin, and R. Zhang, ‘‘Unified low-resource sequence labeling by sample-aware dynamic sparse finetuning,’’ arXiv preprint arXiv:2311.03748, 2023

  116. [123]

    Ansell, E

    A. Ansell, E. M. Ponti, A. Korhonen, and I. Vulić, ‘‘Composable sparse fine-tuning for cross-lingual transfer,’’ arXiv preprint arXiv:2110.07560, 2021

  117. [124]

    Z. Fu, H. Y ang, A. M.-C. So, W. Lam, L. Bing, and N. Collier, ‘‘On the effectiveness of parameter-efficient fine-tuning,’’ in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, 2023, pp. 12 799–12 807

  118. [125]

    R. Xu, F. Luo, Z. Zhang, C. Tan, B. Chang, S. Huang, and F. Huang, ‘‘Raise a child in large language model: Towards effective and generalizable fine-tuning,’’ arXiv preprint arXiv:2109.05687, 2021

  119. [126]

    Vucetic, M

    D. Vucetic, M. Tayaranian, M. Ziaeefard, J. J. Clark, B. H. Meyer, and W. J. Gross, ‘‘Efficient fine-tuning of bert models on the edge,’’ in 2022 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 2022, pp. 1838–1842

  120. [127]

    E. B. Zaken, S. Ravfogel, and Y . Goldberg, ‘‘Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models,’’ arXiv preprint arXiv:2106.10199, 2021

  121. [128]

    Gheini, X

    M. Gheini, X. Ren, and J. May, ‘‘Cross-attention is all you need: Adapting pretrained transformers for machine translation,’’ arXiv preprint arXiv:2104.08771, 2021

  122. [129]

    Aghajanyan, L

    A. Aghajanyan, L. Zettlemoyer, and S. Gupta, ‘‘Intrinsic dimensionality explains the effectiveness of language model fine-tuning,’’ arXiv preprint arXiv:2012.13255, 2020

  123. [131]

    Karimi Mahabadi, J

    R. Karimi Mahabadi, J. Henderson, and S. Ruder, ‘‘Compacter: Efficient low-rank hypercomplex adapter layers,’’Advances in Neural Information Processing Systems, vol. 34, pp. 1022–1035, 2021

  124. [132]

    Edalati, M

    A. Edalati, M. Tahaei, I. Kobyzev, V . P . Nia, J. J. Clark, and M. Rezagholizadeh, ‘‘Krona: Parameter efficient tuning with kronecker adapter,’’arXiv preprint arXiv:2212.10650, 2022

  125. [133]

    X. He, C. Li, P . Zhang, J. Y ang, and X. E. Wang, ‘‘Parameter-efficient model adaptation for vision transformers,’’ in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, 2023, pp. 817–825

  126. [134]

    D. J. Kopiczko, T. Blankevoort, and Y . M. Asano, ‘‘V era: V ector-based random matrix adaptation,’’ arXiv preprint arXiv:2310.11454, 2023

  127. [135]

    Liu, C.-Y

    S.-Y . Liu, C.-Y . Wang, H. Yin, P . Molchanov, Y .-C. F. Wang, K.-T. Cheng, and M.-H. Chen, ‘‘Dora: Weight-decomposed low-rank adaptation,’’ arXiv preprint arXiv:2402.09353, 2024

  128. [136]

    V alipour, M

    M. V alipour, M. Rezagholizadeh, I. Kobyzev, and A. Ghodsi, ‘‘Dylora: Parameter efficient tuning of pre-trained models using dynamic search- free low-rank adaptation,’’ arXiv preprint arXiv:2210.07558, 2022

  129. [137]

    Zhang, M

    Q. Zhang, M. Chen, A. Bukharin, P . He, Y . Cheng, W. Chen, and T. Zhao, ‘‘Adaptive budget allocation for parameter-efficient fine-tuning,’’ arXiv preprint arXiv:2303.10512, 2023

  130. [138]

    N. Ding, X. Lv, Q. Wang, Y . Chen, B. Zhou, Z. Liu, and M. Sun, ‘‘Sparse low-rank adaptation of pre-trained language models,’’ arXiv preprint arXiv:2311.11696, 2023

  131. [139]

    Haobo, H

    S. Haobo, H. Zhao, S. Majumder, and T. Lin, ‘‘Increasing model capacity for free: A simple strategy for parameter efficient fine-tuning,’’ in The Twelfth International Conference on Learning Representations , 2023

  132. [140]

    Zhang, R

    R. Zhang, R. Qiang, S. A. Somayajula, and P . Xie, ‘‘Autolora: Automatically tuning matrix ranks in low-rank adaptation based on meta learning,’’ arXiv preprint arXiv:2403.09113, 2024

  133. [141]

    A. X. Y ang, M. Robeyns, X. Wang, and L. Aitchison, ‘‘Bayesian low-rank adaptation for large language models,’’arXiv preprint arXiv:2308.13111, 2023

  134. [142]

    Y . Lin, X. Ma, X. Chu, Y . Jin, Z. Y ang, Y . Wang, and H. Mei, ‘‘Lora dropout as a sparsity regularizer for overfitting control,’’ arXiv preprint arXiv:2404.09610, 2024

  135. [143]

    X. Meng, D. Dai, W. Luo, Z. Y ang, S. Wu, X. Wang, P . Wang, Q. Dong, L. Chen, and Z. Sui, ‘‘Periodiclora: Breaking the low-rank bottleneck in lora optimization,’’ arXiv preprint arXiv:2402.16141, 2024

  136. [144]

    Hayou, N

    S. Hayou, N. Ghosh, and B. Y u, ‘‘Lora+: Efficient low rank adaptation of large models,’’ arXiv preprint arXiv:2402.12354, 2024

  137. [145]

    Y . Chen, S. Qian, H. Tang, X. Lai, Z. Liu, S. Han, and J. Jia, ‘‘Longlora: Efficient fine-tuning of long-context large language models,’’ 2024. [Online]. Available: https://arxiv.org/abs/2309.12307

  138. [146]

    Huang, Q

    C. Huang, Q. Liu, B. Y . Lin, T. Pang, C. Du, and M. Lin, ‘‘Lorahub: Efficient cross-task generalization via dynamic lora composition,’’ arXiv preprint arXiv:2307.13269, 2023

  139. [147]

    Q. Liu, X. Wu, X. Zhao, Y . Zhu, D. Xu, F. Tian, and Y . Zheng, ‘‘Moelora: An moe-based parameter efficient fine-tuning method for multi-task medical applications,’’ arXiv preprint arXiv:2310.18339, 2023

  140. [148]

    W. Feng, C. Hao, Y . Zhang, Y . Han, and H. Wang, ‘‘Mixture-of-loras: An efficient multitask tuning for large language models,’’ arXiv preprint arXiv:2403.03432, 2024

  141. [149]

    X. Wu, S. Huang, and F. Wei, ‘‘Mixture of lora experts,’’ arXiv preprint arXiv:2404.13628, 2024

  142. [150]

    D. Li, Y . Ma, N. Wang, Z. Cheng, L. Duan, J. Zuo, C. Y ang, and M. Tang, ‘‘Mixlora: Enhancing large language models fine-tuning with lora based mixture of experts,’’ arXiv preprint arXiv:2404.15159, 2024

  143. [151]

    Y . Mao, L. Mathias, R. Hou, A. Almahairi, H. Ma, J. Han, W.-t. Yih, and M. Khabsa, ‘‘Unipelt: A unified framework for parameter-efficient language model tuning,’’ arXiv preprint arXiv:2110.07577, 2021

  144. [152]

    J. Chen, A. Zhang, X. Shi, M. Li, A. Smola, and D. Y ang, ‘‘Parameter-efficient fine-tuning design spaces,’’ arXiv preprint arXiv:2301.01821, 2023

  145. [153]

    Zhang, K

    Y . Zhang, K. Zhou, and Z. Liu, ‘‘Neural prompt search,’’ 2022

  146. [154]

    Zhong, J

    S. Zhong, J. Mo, and Z. Liu, ‘‘Autopet challenge 2022: Automatic segmentation of whole-body tumor lesion based on deep learning and fdg pet/ct,’’ arXiv preprint arXiv:2209.01212, 2022

  147. [155]

    Z. Hu, Y . Lan, L. Wang, W. Xu, E.-P . Lim, R. K.-W. Lee, L. Bing, and S. Poria, ‘‘Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models,’’ arXiv preprint arXiv:2304.01933, 2023

  148. [156]

    S. Hu, Z. Zhang, N. Ding, Y . Wang, Y . Wang, Z. Liu, and M. Sun, ‘‘Sparse structure search for parameter-efficient tuning,’’ arXiv preprint arXiv:2206.07382, 2022. VOLUME 11, 2024 33 Minghao et al.: Survey of different Large Language Model Architectures: Trends, Benchmarks, a...

  149. [157]

    E. J. Hu, Y . Shen, P . Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, ‘‘Lora: Low-rank adaptation of large language models,’’ 2021. [Online]. Available: https://arxiv.org/abs/2106.09685

  150. [158]

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, ‘‘Visual instruction tuning,’’ arXiv preprint arXiv:2304.08485, 2023

  151. [159]

    P . Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P . Clark, and A. Kalyan, ‘‘Learn to explain: Multimodal reasoning via thought chains for science question answering,’’ 2022. [Online]. Available: https://arxiv.org/abs/2209.09513

  152. [160]

    Y .-L. Sung, J. Cho, and M. Bansal, ‘‘Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks,’’ 2022. [Online]. Available: https://arxiv.org/abs/2112.06825

  153. [161]

    Y . Cui, W. Che, T. Liu, B. Qin, and Z. Y ang, ‘‘Pre-training with whole word masking for chinese bert,’’ IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3504–3514, 2021

  154. [162]

    Y . Cui, W. Che, T. Liu, B. Qin, Z. Y ang, S. Wang, and G. Hu, ‘‘Pre-training with whole word masking for chinese bert,’’ arXiv preprint arXiv:1906.08101, 2019

  155. [163]

    Joshi, D

    M. Joshi, D. Chen, Y . Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, ‘‘Spanbert: Improving pre-training by representing and predicting spans,’’ Transactions of the association for computational linguistics , vol. 8, pp. 64–77, 2020

  156. [164]

    V . Sanh, L. Debut, J. Chaumond, and T. Wolf, ‘‘Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,’’ arXiv preprint arXiv:1910.01108, 2019

  157. [165]

    Y . Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V . Stoyanov, ‘‘Roberta: A robustly optimized bert pretraining approach,’’ arXiv preprint arXiv:1907.11692, 2019

  158. [166]

    X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, ‘‘Tinybert: Distilling bert for natural language understanding,’’ arXiv preprint arXiv:1909.10351, 2019

  159. [167]

    L. H. Li, M. Y atskar, D. Yin, C.-J. Hsieh, and K.-W. Chang, ‘‘Visualbert: A simple and performant baseline for vision and language,’’ arXiv preprint arXiv:1908.03557, 2019

  160. [168]

    Y . Cui, W. Che, T. Liu, B. Qin, S. Wang, and G. Hu, ‘‘Revisiting pre-trained models for chinese natural language processing,’’ arXiv preprint arXiv:2004.13922, 2020

  161. [169]

    H. Bao, L. Dong, S. Piao, and F. Wei, ‘‘Beit: Bert pre-training of image transformers,’’ arXiv preprint arXiv:2106.08254, 2021

  162. [170]

    Z. Peng, L. Dong, H. Bao, Q. Y e, and F. Wei, ‘‘Beit v2: Masked image modeling with vector-quantized visual tokenizers,’’ arXiv preprint arXiv:2208.06366, 2022

  163. [171]

    W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som et al. , ‘‘Image as a foreign language: Beit pretraining for all vision and vision-language tasks,’’ arXiv preprint arXiv:2208.10442, 2022

  164. [172]

    V ahdat, E

    A. V ahdat, E. Andriyash, and W. Macready, ‘‘Dvae#: Discrete variational autoencoders with relaxed boltzmann priors,’’ Advances in Neural Information Processing Systems, vol. 31, 2018

  165. [173]

    Xu, ‘‘Roberta-wwm-ext fine-tuning for chinese text classification,’’ arXiv preprint arXiv:2103.00492, 2021

    Z. Xu, ‘‘Roberta-wwm-ext fine-tuning for chinese text classification,’’ arXiv preprint arXiv:2103.00492, 2021

  166. [174]

    Y . Sun, S. Wang, Y . Li, S. Feng, H. Tian, H. Wu, and H. Wang, ‘‘Ernie 2.0: A continual pre-training framework for language understanding,’’ in Proceedings of the AAAI conference on artificial intelligence , vol. 34, 2020, pp. 8968–8975

  167. [175]

    Y . Sun, S. Wang, S. Feng, S. Ding, C. Pang, J. Shang, J. Liu, X. Chen, Y . Zhao, Y . Lu et al. , ‘‘Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation,’’ arXiv preprint arXiv:2107.02137, 2021

  168. [176]

    F. Y u, J. Tang, W. Yin, Y . Sun, H. Tian, H. Wu, and H. Wang, ‘‘Ernie-vil: Knowledge enhanced vision-language representations through scene graphs,’’ inProceedings of the AAAI Conference on Artificial Intelligence, vol. 35, 2021, pp. 3208–3216

  169. [177]

    B. Shan, W. Yin, Y . Sun, H. Tian, H. Wu, and H. Wang, ‘‘Ernie-vil 2.0: Multi-view contrastive learning for image-text pre-training,’’ arXiv preprint arXiv:2209.15270, 2022

  170. [178]

    Clark, M.-T

    K. Clark, M.-T. Luong, Q. V . Le, and C. D. Manning, ‘‘Electra: Pre-training text encoders as discriminators rather than generators,’’ arXiv preprint arXiv:2003.10555, 2020

  171. [179]

    Goodfellow, J

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, ‘‘Generative adversarial networks,’’ Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020

  172. [180]

    P . He, X. Liu, J. Gao, and W. Chen, ‘‘Deberta: Decoding-enhanced bert with disentangled attention,’’ arXiv preprint arXiv:2006.03654, 2020

  173. [181]

    Z. Dai, Z. Y ang, Y . Y ang, J. Carbonell, Q. V . Le, and R. Salakhutdinov, ‘‘Transformer-xl: Attentive language models beyond a fixed-length context,’’arXiv preprint arXiv:1901.02860, 2019

  174. [182]

    Y ang, Z

    Z. Y ang, Z. Dai, Y . Y ang, J. Carbonell, R. R. Salakhutdinov, and Q. V . Le, ‘‘Xlnet: Generalized autoregressive pretraining for language understand- ing,’’ Advances in neural information processing systems , vol. 32, 2019

  175. [183]

    Y . Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, ‘‘Aligning books and movies: Towards story-like visual explanations by watching movies and reading books,’’ in The IEEE International Conference on Computer Vision (ICCV) , December 2015

  176. [184]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , ‘‘Language models are unsupervised multitask learners,’’ OpenAI blog , vol. 1, no. 8, p. 9, 2019

  177. [185]

    McCann, N

    B. McCann, N. S. Keskar, C. Xiong, and R. Socher, ‘‘The natural language decathlon: Multitask learning as question answering,’’ arXiv preprint arXiv:1806.08730, 2018

  178. [186]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P . Dhariwal, A. Neelakantan, P . Shyam, G. Sastry, A. Askellet al., ‘‘Language models are few-shot learners,’’ Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020

  179. [187]

    C. Finn, P . Abbeel, and S. Levine, ‘‘Model-agnostic meta-learning for fast adaptation of deep networks,’’ in International conference on machine learning. PMLR, 2017, pp. 1126–1135

  180. [188]

    [Online]

    ‘‘commoncrawl,’’ 2022. [Online]. Available: https://commoncrawl.org/

  181. [189]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , ‘‘Training language models to follow instructions with human feedback,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 27 730–27 744, 2022

  182. [190]

    [Online]

    OpenAI, ‘‘Chatgpt,’’ 2022. [Online]. Available: https://openai.com/chatgpt

  183. [191]

    ——, ‘‘Gpt-4 technical report,’’ 2023

  184. [192]

    Black, L

    S. Black, L. Gao, P . Wang, C. Leahy, and S. R. Biderman, ‘‘Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow,’’ 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:245758737

  185. [193]

    Z. Hu, Y . Dong, K. Wang, K.-W. Chang, and Y . Sun, ‘‘Gpt-gnn: Generative pre-training of graph neural networks,’’ in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020, pp. 1857–1867

  186. [194]

    Komatsuzaki, ‘‘Gpt-j-6b: 6b jax-based transformer,’’ Jun 2021

    A. Komatsuzaki, ‘‘Gpt-j-6b: 6b jax-based transformer,’’ Jun 2021. [Online]. Available: https://arankomatsuzaki.wordpress.com/2021/06/ 04/gpt-j/

  187. [195]

    Huebner, ‘‘Mesh transformer jax,’’ Feb 2023

    C. Huebner, ‘‘Mesh transformer jax,’’ Feb 2023. [Online]. Available: https://www.eleuther.ai/artifacts/mtj

  188. [196]

    Black, S

    S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang et al. , ‘‘Gpt-neox-20b: An open-source autoregressive language model,’’ arXiv preprint arXiv:2204.06745, 2022

  189. [197]

    Microsoft, ‘‘Deepspeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.’’

  190. [198]

    Zhang, S

    Y . Zhang, S. Sun, M. Galley, Y .-C. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, and B. Dolan, ‘‘Dialogpt: Large-scale generative pre-training for conversational response generation,’’ arXiv preprint arXiv:1911.00536 , 2019

  191. [199]

    Google, ‘‘Introducing pathways: A next-generation ai architecture,’’

  192. [200]

    ——, ‘‘Pathways language model (palm): Scaling to 540 billion parameters for breakthrough performance,’’

  193. [201]

    Available: https://blog.google/technology/ai/ introducing-pathways-next-generation-ai-architecture/

    [Online]. Available: https://blog.google/technology/ai/ introducing-pathways-next-generation-ai-architecture/

  194. [202]

    J. Su, Y . Lu, S. Pan, A. Murtadha, B. Wen, and Y . Liu, ‘‘Roformer: Enhanced transformer with rotary position embedding,’’ arXiv preprint arXiv:2104.09864, 2021

  195. [203]

    Available: https://blog.research.google/2022/04/ pathways-language-model-palm-scaling-to.html

    [Online]. Available: https://blog.research.google/2022/04/ pathways-language-model-palm-scaling-to.html

  196. [204]

    Shazeer, ‘‘Glu variants improve transformer,’’ arXiv preprint arXiv:2002.05202, 2020

    N. Shazeer, ‘‘Glu variants improve transformer,’’ arXiv preprint arXiv:2002.05202, 2020

  197. [205]

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P . Perona, D. Ramanan, P . Dollár, and C. L. Zitnick, ‘‘Microsoft coco: Common objects in context,’’ in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springe...

  198. [206]

    Driess, F

    D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Y u et al. , ‘‘Palm-e: An embodied multimodal language model,’’ arXiv preprint arXiv:2303.03378, 2023

  199. [207]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ‘‘Imagenet: A large-scale hierarchical image database,’’ in 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255. 34 VOLUME 11, 2024 Minghao et al.: Survey of different Large Language M...

  200. [208]

    Mehta, A

    H. Mehta, A. Thakurta, A. Kurakin, and A. Cutkosky, ‘‘Large scale transfer learning for differentially private image classification,’’ arXiv preprint arXiv:2205.02973, 2022

  201. [209]

    Krishna, Y

    R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L.-J. Li, D. A. Shammaet al., ‘‘Visual genome: Connecting language and vision using crowdsourced dense image annotations,’’ International journal of computer vision , vol. 123, pp. 32–73, 2017

  202. [210]

    Dehghani, J

    M. Dehghani, J. Djolonga, B. Mustafa, P . Padlewski, J. Heek, J. Gilmer, A. P . Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin et al. , ‘‘Scaling vision transformers to 22 billion parameters,’’ inInternational Conference on Machine Learning. PMLR, 2023, pp. 7480–7512

  203. [211]

    X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P . Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer et al. , ‘‘Pali: A jointly-scaled multilingual language-image model,’’ arXiv preprint arXiv:2209.06794, 2022

  204. [212]

    Huang, L

    S. Huang, L. Dong, W. Wang, Y . Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, Q. Liu et al., ‘‘Language is not all you need: Aligning perception with language models,’’ arXiv preprint arXiv:2302.14045 , 2023

  205. [213]

    Lewkowycz, A

    A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Soloet al., ‘‘Solving quantitative reasoning problems with language models,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 3843–3857, 2022

  206. [214]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, and T. Unterthiner, ‘‘Transformers for image recognition at scale,’’ arXiv preprint arXiv:2010.11929, 2020

  207. [215]

    L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, ‘‘mt5: A massively multilingual pre-trained text-to-text transformer,’’arXiv preprint arXiv:2010.11934, 2020

  208. [216]

    X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, ‘‘Scaling vision transformers,’’ inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12 104–12 113

  209. [217]

    D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P . Szolovits, ‘‘What disease does this patient have? a large-scale open domain question answering dataset from medical exams,’’ Applied Sciences , vol. 11, no. 14, p. 6421, 2021

  210. [218]

    R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P . Bailey, Z. Chenet al., ‘‘Palm 2 technical report,’’ arXiv preprint arXiv:2305.10403, 2023

  211. [219]

    Singhal, T

    K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, L. Hou, K. Clark, S. Pfohl, H. Cole-Lewis, D. Neal et al. , ‘‘Towards expert-level medical question answering with large language models,’’ arXiv preprint arXiv:2305.09617, 2023

  212. [221]

    Singhal, S

    K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl et al. , ‘‘Large language models encode clinical knowledge,’’ arXiv preprint arXiv:2212.13138, 2022

  213. [222]

    H. Wang, S. Ma, S. Huang, L. Dong, W. Wang, Z. Peng, Y . Wu, P . Bajaj, S. Singhal, A. Benhaim et al., ‘‘Foundation transformers,’’arXiv preprint arXiv:2210.06423, 2022

  214. [223]

    V . A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, ‘‘Reducing activation recomputation in large transformer models,’’ Proceedings of Machine Learning and Systems, vol. 5, 2023

  215. [224]

    Shoeybi, M

    M. Shoeybi, M. Patwary, R. Puri, P . LeGresley, J. Casper, and B. Catan- zaro, ‘‘Megatron-lm: Training multi-billion parameter language models using model parallelism,’’ arXiv preprint arXiv:1909.08053, 2019

  216. [225]

    Narayanan, M

    D. Narayanan, M. Shoeybi, J. Casper, P . LeGresley, M. Patwary, V . Korthikanti, D. V ainbrand, P . Kashinkunti, J. Bernauer, B. Catanzaro et al. , ‘‘Efficient large-scale language model training on gpu clusters using megatron-lm,’’ in Proceedings of the International Conferen...

  217. [226]

    Smith, M

    S. Smith, M. Patwary, B. Norick, P . LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V . Korthikanti et al. , ‘‘Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,’’ arXiv preprint arXiv:2201.11990, 2022

  218. [227]

    ‘‘Turing-nlg: A 17-billion-parameter language model by microsoft,’’ Feb

  219. [228]

    P . Xu, M. Patwary, M. Shoeybi, R. Puri, P . Fung, A. Anandkumar, and B. Catanzaro, ‘‘Megatron-cntrl: Controllable story generation with external knowledge using large-scale language models,’’ arXiv preprint arXiv:2010.00840, 2020

  220. [229]

    Rajbhandari, J

    S. Rajbhandari, J. Rasley, O. Ruwase, and Y . He, ‘‘Zero: Memory optimizations toward training trillion parameter models,’’ in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2020, pp. 1–16

  221. [230]

    Taori, I

    R. Taori, I. Gulrajani, T. Zhang, Y . Dubois, X. Li, C. Guestrin, P . Liang, and T. B. Hashimoto, ‘‘Alpaca: A strong, replicable instruction-following model,’’ 2019

  222. [231]

    H.-C. Shin, Y . Zhang, E. Bakhturina, R. Puri, M. Patwary, M. Shoeybi, and R. Mani, ‘‘Biomegatron: Larger biomedical domain language model,’’ arXiv preprint arXiv:2010.06060, 2020

  223. [232]

    [Online]

    ‘‘Guanaco - generative universal assistant for natural-language adaptive context-aware omnilingual outputs,’’ Feb 2020. [Online]. Available: https://guanaco-model.github.io/

  224. [233]

    Zhang and R

    B. Zhang and R. Sennrich, ‘‘Root mean square layer normalization,’’ Advances in Neural Information Processing Systems , vol. 32, 2019

  225. [234]

    Chiang, Z

    W.-L. Chiang, Z. Li, Z. Lin, Y . Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y . Zhuang, J. E. Gonzalez, I. Stoica, and E. P . Xing, ‘‘Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,’’ March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-...

  226. [235]

    [Online]

    OpenAI, ‘‘Models - openai,’’ Feb 2020. [Online]. Available: https://platform.openai.com/docs/models

  227. [236]

    [Online]

    Databricks, ‘‘dolly,’’ Feb 2020. [Online]. Available: https://github.com/databrickslabs/dolly

  228. [237]

    [Online]

    ‘‘Alpaca-lora,’’ Feb 2020. [Online]. Available: https: //github.com/tloen/alpaca-lora

  229. [238]

    Biderman, H

    S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff et al. , ‘‘Pythia: A suite for analyzing large language models across training and scaling,’’ in International Conference on Machine Learning . PMLR...

  230. [239]

    [Online]

    ShareGPT, ‘‘Sharegpt,’’ Feb 2020. [Online]. Available: https://sharegpt.com/

  231. [240]

    Zheng, W.-L

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P . Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, ‘‘Judging llm-as-a-judge with mt-bench and chatbot arena,’’ 2023

  232. [241]

    Conover, M

    M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P . Wendell, M. Zaharia, and R. Xin, ‘‘Free dolly: Introducing the world’s first truly open instruction-tuned llm,’’

  233. [242]

    L. L. Ziang Leng, Qiyuan Chen, ‘‘Luotuo: Chinese-alpaca-lora,’’ March 2023. [Online]. Available: https://github.com/LC1332/Chinese-alpaca-lora

  234. [243]

    Y . Cui, Z. Y ang, and X. Y ao, ‘‘Efficient and effective text encoding for chinese llama and alpaca,’’ arXiv preprint arXiv:2304.08177 , 2023. [Online]. Available: https://arxiv.org/abs/2304.08177

  235. [244]

    X. Geng, A. Gudibande, H. Liu, E. Wallace, P . Abbeel, S. Levine, and D. Song, ‘‘Koala: A dialogue model for academic research,’’ Apr 2023. [Online]. Available: https://bair.berkeley.edu/blog/2023/04/03/koala/

  236. [245]

    P . Gao, J. Han, R. Zhang, Z. Lin, S. Geng, A. Zhou, W. Zhang, P . Lu, C. He, X. Y ue et al. , ‘‘Llama-adapter v2: Parameter-efficient visual instruction model,’’ arXiv preprint arXiv:2304.15010, 2023

  237. [246]

    C. Xu, D. Guo, N. Duan, and J. McAuley, ‘‘Baize: An open-source chat model with parameter-efficient tuning on self-chat data,’’ arXiv preprint arXiv:2304.01196, 2023

  238. [247]

    LLaV A, ‘‘Chinese llava,’’ July 2023

    C. LLaV A, ‘‘Chinese llava,’’ July 2023. [Online]. Available: https://github.com/LinkSoul-AI/Chinese-LLaV A VOLUME 11, 2024 35 Minghao et al.: Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges

  239. [248]

    Y . Shu, S. Dong, G. Chen, W. Huang, R. Zhang, D. Shi, Q. Xiang, and Y . Shi, ‘‘Llasm: Large language and speech model,’’ arXiv preprint arXiv:2308.15930, 2023

  240. [249]

    Zhang, J

    R. Zhang, J. Han, A. Zhou, X. Hu, S. Y an, P . Lu, H. Li, P . Gao, and Y . Qiao, ‘‘Llama-adapter: Efficient fine-tuning of language models with zero-init attention,’’ arXiv preprint arXiv:2303.16199, 2023

  241. [250]

    Zhang, X

    H. Zhang, X. Li, and L. Bing, ‘‘Video-llama: An instruction-tuned audio-visual language model for video understanding,’’ arXiv preprint arXiv:2306.02858, 2023

  242. [251]

    D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, ‘‘Minigpt-4: Enhancing vision-language understanding with advanced large language models,’’ arXiv preprint arXiv:2304.10592, 2023

  243. [252]

    K. Li, Y . He, Y . Wang, Y . Li, W. Wang, P . Luo, Y . Wang, L. Wang, and Y . Qiao, ‘‘Videochat: Chat-centric video understanding,’’ arXiv preprint arXiv:2305.06355, 2023

  244. [253]

    Girdhar, A

    R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, ‘‘Imagebind: One embedding space to bind them all,’’ in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 15 180–15 190

  245. [254]

    [Online]

    Visual-LLaMA, ‘‘Visual-llama,’’ 2023. [Online]. Available: https://github.com/feizc/Visual-LLaMA

  246. [255]

    Kudo and J

    T. Kudo and J. Richardson, ‘‘Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,’’ arXiv preprint arXiv:1808.06226, 2018

  247. [256]

    Zhang, J

    Q. Zhang, J. Zhang, Y . Xu, and D. Tao, ‘‘Vision transformer with quadrangle attention,’’ arXiv preprint arXiv:2303.15105, 2023

  248. [257]

    Hennigan, T

    T. Hennigan, T. Cai, T. Norman, L. Martens, and I. Babuschkin, ‘‘Haiku: Sonnet for JAX,’’ 2020. [Online]. Available: http://github.com/deepmind/dm-haiku

  249. [258]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al., ‘‘Training compute-optimal large language models,’’ arXiv preprint arXiv:2203.15556, 2022

  250. [259]

    J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Y oung et al., ‘‘Scaling language models: Methods, analysis & insights from training gopher,’’ arXiv preprint arXiv:2112.11446, 2021

  251. [260]

    Brock, S

    A. Brock, S. De, S. L. Smith, and K. Simonyan, ‘‘High-performance large-scale image recognition without normalization,’’ in International Conference on Machine Learning. PMLR, 2021, pp. 1059–1071

  252. [261]

    Bradbury, R

    J. Bradbury, R. Frostig, P . Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. V anderPlas, S. Wanderman-Milne, and Q. Zhang, ‘‘JAX: composable transformations of Python+NumPy programs,’’ 2018. [Online]. Available: http://github.com/google/jax

  253. [262]

    Lieber, O

    O. Lieber, O. Sharir, B. Lenz, and Y . Shoham, ‘‘Jurassic-1: Technical details and evaluation,’’ White Paper . AI21 Labs, vol. 1, 2021

  254. [263]

    Karpas, O

    E. Karpas, O. Abend, Y . Belinkov, B. Lenz, O. Lieber, N. Ratner, Y . Shoham, H. Bata, Y . Levine, K. Leyton-Brownet al., ‘‘Mrkl systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning,’’ arXiv prep...

  255. [264]

    Alayrac, J

    J.-B. Alayrac, J. Donahue, P . Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , ‘‘Flamingo: a visual language model for few-shot learning,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 23 716–23 736, 2022

  256. [265]

    [Online]

    Anthropic, ‘‘Introducing claude,’’ Mar 2023. [Online]. Available: https://www.anthropic.com/index/introducing-claude

  257. [266]

    Lab, ‘‘Ai21 studio logo,’’ August 2021

    A. Lab, ‘‘Ai21 studio logo,’’ August 2021. [Online]. Available: https://www.ai21.com/studio

  258. [267]

    L. Mei, J. Mao, Z. Wang, C. Gan, and J. B. Tenenbaum, ‘‘Falcon: fast visual concept learning by integrating images, linguistic descriptions, and conceptual relations,’’ arXiv preprint arXiv:2203.16639, 2022

  259. [268]

    Penedo, Q

    G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay, ‘‘The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only,’’ arXiv preprint arXiv:2306.01116, 2023

  260. [269]

    Lab, ‘‘Announcing jurassic-2 and task-specific apis,’’ March 2022

    A. Lab, ‘‘Announcing jurassic-2 and task-specific apis,’’ March 2022. [Online]. Available: https://www.ai21.com/blog/introducing-j2

  261. [270]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P . Mishkin, J. Clark et al. , ‘‘Learning transferable visual models from natural language supervision,’’ in International conference on machine learning. PMLR, 2021, pp. 8748–8763

  262. [271]

    [Online]

    ——, ‘‘Claude 2,’’ Jul 2023. [Online]. Available: https://www.anthropic.com/index/claude-2

  263. [272]

    [Online]

    OpenAI, ‘‘Dall ·e 3,’’ September 2023. [Online]. Available: https://openai.com/dall-e-3

  264. [273]

    Radford, J

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, ‘‘Robust speech recognition via large-scale weak supervision,’’ in International Conference on Machine Learning . PMLR, 2023, pp. 28 492–28 518

  265. [274]

    Ramesh, M

    A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. V oss, A. Radford, M. Chen, and I. Sutskever, ‘‘Zero-shot text-to-image generation,’’ in International Conference on Machine Learning. PMLR, 2021, pp. 8821–8831

  266. [275]

    R. Li, L. B. Allal, Y . Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim et al. , ‘‘Starcoder: may the source be with you!’’ arXiv preprint arXiv:2305.06161, 2023

  267. [276]

    Ramesh, P

    A. Ramesh, P . Dhariwal, A. Nichol, C. Chu, and M. Chen, ‘‘Hierarchical text-conditional image generation with clip latents,’’ arXiv preprint arXiv:2204.06125, vol. 1, no. 2, p. 3, 2022

  268. [277]

    Thoppilan, D

    R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y . Du et al., ‘‘Lamda: Language models for dialog applications,’’ arXiv preprint arXiv:2201.08239, 2022

  269. [278]

    G. H. Cohen, ‘‘Align: a program to superimpose protein coordinates, accounting for insertions and deletions,’’ Journal of applied crystallography, vol. 30, no. 6, pp. 1160–1161, 1997

  270. [279]

    M. Chen, J. Tworek, H. Jun, Q. Y uan, H. P . d. O. Pinto, J. Kaplan, H. Ed- wards, Y . Burda, N. Joseph, G. Brockman et al. , ‘‘Evaluating large lan- guage models trained on code,’’ arXiv preprint arXiv:2107.03374, 2021

  271. [280]

    N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Y u, O. Firatet al., ‘‘Glam: Efficient scaling of language models with mixture-of-experts,’’ in International Conference on Machine Learning. PMLR, 2022, pp. 5547–5569

  272. [281]

    D. So, Q. Le, and C. Liang, ‘‘The evolved transformer,’’ in International conference on machine learning. PMLR, 2019, pp. 5877–5886

  273. [282]

    [Online]

    ‘‘Phi-1.5,’’ 2023. [Online]. Available: https://huggingface.co/microsoft/ phi-1_5

  274. [283]

    [Online]

    ‘‘Phi-2: The surprising power of small language models,’’ 2023. [Online]. Available: https://www.microsoft.com/en-us/research/blog/ phi-2-the-surprising-power-of-small-language-models/

  275. [284]

    Koonce and B

    B. Koonce and B. Koonce, ‘‘Efficientnet,’’ Convolutional Neural Networks with Swift for Tensorflow: Image Recognition and Dataset Categorization, pp. 109–123, 2021

  276. [285]

    H. Xu, Q. Y e, M. Y an, Y . Shi, J. Y e, Y . Xu, C. Li, B. Bi, Q. Qian, W. Wang et al. , ‘‘mplug-2: A modularized multi-modal foundation model across text, image and video,’’ arXiv preprint arXiv:2302.00402, 2023

  277. [286]

    G. Team, R. Anil, S. Borgeaud, Y . Wu, J.-B. Alayrac, J. Y u, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth et al. , ‘‘Gemini: a family of highly capable multimodal models,’’ arXiv preprint arXiv:2312.11805, 2023

  278. [287]

    Y . Shen, K. Song, X. Tan, D. Li, W. Lu, and Y . Zhuang, ‘‘Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface,’’ arXiv preprint arXiv:2303.17580, 2023

  279. [288]

    H. Xu, Q. Y e, X. Wu, M. Y an, Y . Miao, J. Y e, G. Xu, A. Hu, Y . Shi, G. Xu et al. , ‘‘Y ouku-mplug: A 10 million large-scale chinese video-language dataset for pre-training and benchmarks,’’ arXiv preprint arXiv:2306.04362, 2023

  280. [289]

    C. Li, H. Xu, J. Tian, W. Wang, M. Y an, B. Bi, J. Y e, H. Chen, G. Xu, Z. Cao et al., ‘‘mplug: Effective and efficient vision-language learning by cross-modal skip-connections,’’ arXiv preprint arXiv:2205.12005, 2022

  281. [290]

    Sanders, D

    K. Sanders, D. Etter, R. Kriz, and B. V an Durme, ‘‘Multivent: Multilingual videos of events with aligned natural text,’’ arXiv preprint arXiv:2307.03153, 2023

  282. [291]

    Y ang, L

    Z. Y ang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, ‘‘Mm-react: Prompting chatgpt for multimodal reasoning and action,’’ arXiv preprint arXiv:2303.11381, 2023

  283. [292]

    S. Bao, H. He, F. Wang, H. Wu, and H. Wang, ‘‘Plato: Pre-trained dialogue generation model with discrete latent variable,’’ arXiv preprint arXiv:1910.07931, 2019

  284. [293]

    S. Bao, H. He, F. Wang, H. Wu, H. Wang, W. Wu, Z. Guo, Z. Liu, and X. Xu, ‘‘Plato-2: Towards building an open-domain chatbot via curriculum learning,’’ arXiv preprint arXiv:2006.16779, 2020

  285. [294]

    J. Y e, A. Hu, H. Xu, Q. Y e, M. Y an, Y . Dan, C. Zhao, G. Xu, C. Li, J. Tian et al. , ‘‘mplug-docowl: Modularized multimodal large language model for document understanding,’’ arXiv preprint arXiv:2307.02499, 2023

  286. [295]

    Y . Huo, M. Zhang, G. Liu, H. Lu, Y . Gao, G. Y ang, J. Wen, H. Zhang, B. Xu, W. Zheng et al., ‘‘Wenlan: Bridging vision and language by large- scale multi-modal pre-training,’’ arXiv preprint arXiv:2103.06561, 2021

  287. [296]

    Soltan, S

    S. Soltan, S. Ananthakrishnan, J. FitzGerald, R. Gupta, W. Hamza, H. Khan, C. Peris, S. Rawls, A. Rosenbaum, A. Rumshisky et al. , ‘‘Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model,’’ arXiv preprint arXiv:2208.01448, 2022

  288. [299]

    [Online]

    BAAI, ‘‘Baai 23,’’ 2023. [Online]. Available: https://2023.baai.ac.cn/about

  289. [2020]

    Available: https://www.microsoft.com/en-us/research/ blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/

    [Online]. Available: https://www.microsoft.com/en-us/research/ blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/

  290. [2022]

    Available: https://github.com/microsoft/DeepSpeed

    [Online]. Available: https://github.com/microsoft/DeepSpeed

  291. [2023]

    Available: https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially-viable-instruction-tuned-llm

    [Online]. Available: https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially-viable-instruction-tuned-llm

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.