REVIEW 3 major objections 4 minor 78 references
A Primer on Large Language Models and their Limitations
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read LLMs do not hallucinate; they produce bullshit in the technical sense, the authors argue.
desk verdict A competent and occasionally useful LLM primer that trips over its own philosophical headline: the claim that LLMs 'produce bullshit' is undercut by the authors' own verifiable example. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the concept of bullshit in its technical sense: communication produced with indifference to the truth. The mechanism that earns this label is the decoder's next-token prediction objective, which optimizes for plausible continuation rather than factual correctness. The paper additionally relies on the premise that no algorithm for truth exists, so judging an output to be a hallucination is a subjective human evaluation. This framing does the work of shifting the mitigation question from 'how do we fix the model's perception?' to 'how do we manage a generator that is structurally unconcerned with truth?'
What would settle it
A single reproducible case in which a model's generation demonstrably depends on the truth of the proposition—for example, a model that consistently refuses to state a false claim even when the prompt rewards falsehood—would show that indifference to truth is not total and would undercut the blanket claim that all LLM output is bullshit.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that LLM errors belong to a different category than perceptual errors. Because decoder-based language models are designed to produce the most plausible continuation of a prompt, they are indifferent to the factual status of what they generate; they do not first form a belief and then misstate it. Calling an output a 'hallucination' treats the error as a failure of perception, but the authors argue this is a post-hoc value judgment made by a human, since there is no algorithm for truth against which the output can be checked. The appropriate description, they argue, is bullshit in the technical sense: speech made without concern for the truth, which may accidentally be true or false. The paper supports this by noting that several different LLMs generated similar false references when asked for canonical academic citations, which it attributes to shared architectures and training data.
Load-bearing premise
The argument depends on the premise that there is no algorithm for truth, so calling any output a hallucination is a subjective human judgment rather than a factual description.
Editorial extensions
If this is right
- If LLM errors are truth-indifferent rather than perceptual, then no amount of scaling or fine-tuning will make a hallucination-free model; the goal becomes detecting and managing untrustworthy output.
- Users should treat LLM output as a starting point, verifying critical claims against trusted sources instead of relying on the model's fluency.
- The term 'hallucination' should be retired in technical discussion, because it smuggles in a subjective judgment that the output is false rather than describing the generation process.
- Similar false outputs across different LLMs are to be expected, since models share architectures and training data, so cross-model agreement is not evidence of correctness.
- Mitigation strategies for related risks—jailbreak attacks, catastrophic forgetting, model collapse—are partial, so continual evaluation of model outputs remains necessary.
Reading between the lines
- The paper does not spell this out, but the bullshit framing implies that verification should be built into LLM applications: retrieval-augmented generation and tool use are not just enhancements but the primary defense against truth-indifferent generation.
- If the 'no algorithm for truth' premise is taken literally, it also rules out any objective benchmark of factual accuracy, which would undercut the very hallucination indexes the paper cites; a more moderate reading would restrict the claim to open-ended generation.
- A testable extension would be to measure whether models can be trained to emit calibrated uncertainty or explicit 'I don't know' responses; if such training succeeds, it would show that indifference to truth can be reduced even if not eliminated.
- The bullshit framing invites a philosophical worry the paper sets aside: bullshit originally describes an intentional human attitude, and applying it to a statistical model may be metaphorical; the practical conclusions survive that worry, but the name may not.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a primer on large language models: it surveys transformer architectures (encoder-only, decoder-only, and encoder-decoder), pre-training objectives, fine-tuning and prompt-engineering techniques, orchestration with retrieval and knowledge-representation systems, and a set of risks (catastrophic forgetting, model collapse, jailbreaks, and hallucination). Its most distinctive and emphasized claim is in Section 4.4: LLMs do not hallucinate, they produce bullshit in Frankfurt's technical sense, because they generate with indifference to truth; the paper argues that the term 'hallucination' is a post-hoc human value judgment and that one cannot separate true from false outputs because 'there is no algorithm for truth.' The paper is written for a broad academic and industry audience and relies entirely on cited literature plus a few anecdotal own-experiments.
Significance. As a synthesis, the survey portions have real value for non-specialists: the organizational structure is clear, the figures and tables are helpful, and the treatment of orchestration and mitigation strategies is pragmatic. The paper is also honest about limitations and cites recent, relevant work. Its conceptual contribution is concentrated in Section 4.4, where the 'bullshit' reframing borrows directly from Hannigan et al. and Hicks et al. but adds a disputed premise about the absence of an 'algorithm for truth.' That premise is not defended and is contradicted by the paper's own citation-checking example, so the paper's most distinctive claim currently rests on unstable ground. The factual misattributions in the survey further reduce confidence in a document whose stated purpose is to orient newcomers.
major comments (3)
- [Section 4.4, 'Hallucinations and their Impacts'] The claim that 'one cannot separate one language generation output from another in a meaningful way because there is no algorithm for truth' is overstated and internally contradicted by the authors' own verification example in the same section. The authors report that suggested academic references 'did not exist,' that 'links to the papers did not resolve,' and that 'conferences and journals cited do not list the papers.' These are precisely algorithmic or quasi-algorithmic checks of factual adequacy. General undecidability of truth (e.g., for arbitrary mathematical statements) does not imply that all factual claims are inseparable; '2+2=4' versus '2+2=5,' or an existing DOI versus a fabricated one, are decidably separable. Since this premise is load-bearing for the argument that 'hallucination' is merely a subjective judgment, the reframing is unsupported as written.
- [Section 4.4, Frankfurtian bullshit definition] The paper adopts Frankfurt's definition of bullshit as communication by a speaker who is 'indifferent to the truth,' but it does not supply an operational criterion for what indifference means for a system with no beliefs, desires, or intentions. The citation to Hicks et al. does not resolve this difficulty because the paper's own formulation ties the bullshit conclusion to the 'no algorithm for truth' premise, which fails as argued above. Without either a defended account of model-level indifference or a revised definition, the conclusion that 'LLMs do not hallucinate, they produce bullshit' does not follow from the premises given.
- [Sections 2.2 and 2.6] The survey contains several factual errors that are significant for a primer whose purpose is reliable orientation: RoBERTa is attributed to 'researchers at Google' (Section 2.2), 'peer-to-peer (p2p) from Google' is listed as a decoder-only model (Section 2.2.1), and the GPT listing in Section 2 includes 'GPT-1o preview and GPT-1o mini,' which are inconsistent with the actual model names. These errors are easily corrected but, in a document aimed at non-specialists, they materially undermine trust in the survey's accuracy. They should be fixed before publication.
minor comments (4)
- [Throughout] Typos and stylistic slips: 'property view' should be 'properly view' (Section 4.4), 'in tact' should be 'intact' (Section 2.5), 'vasts amounts' should be 'vast amounts' (Section 2.3), 'accomodate' should be 'accommodate' (Section 2.1), and 'lastlycompletion' should be 'lastly, completion' (Section 2.6.1).
- [Section 2.6] The paper inconsistently names models and products, e.g., 'BARD' instead of 'Bard,' and the list of GPT variants (GPT-3, GPT-4, GPT-4o, GPT-1o preview, GPT-1o mini) is internally inconsistent. Please align names with the official product names.
- [Section 2.5] The statement that the training data allocation is 'usually set around 15%' appears without a citation and is not generally true; typical splits are larger. Please clarify or remove.
- [Section 4.5] The claim that ChatGPT o1-preview 'will correctly answer 3 because it does parse the word and then double check itself' is presented as a definitive fact without a systematic test; given the paper's own emphasis on anecdotal evidence, this should be softened or supported.
Circularity Check
No circularity found: the paper is a survey-primer whose claims are grounded in external sources, and its central 'bullshit' argument transparently imports a defined term from cited philosophical literature rather than deriving it from its own conclusions.
full rationale
This manuscript does not contain a derivation chain in the sense of equations, fitted parameters, or predictions built from its own outputs. The central conceptual claim in Section 4.4, that LLMs 'do not hallucinate, they produce bullshit, in that word's technical sense,' is explicitly imported from Frankfurt's 'On Bullshit' via citations to Hannigan et al. and Hicks et al., with the authors quoting Hicks directly and attributing 'botshit' to Hannigan et al. This is transparent definitional borrowing from external work, not a self-referential reduction. The supporting observation about non-existent references is anecdotal evidence, not a claim derived from the conclusion. The paper's other sections summarize prior literature (e.g., Transformer architecture from Vaswani et al., fine-tuning methods from cited surveys, hallucination indexes from Galileo) without making predictions that reduce to their inputs. There are no self-citations by Johnson and Hyland-Wood that carry argumentative weight, no fitted parameter presented as a prediction, and no uniqueness theorem imported from the authors' own prior work. The philosophical premise that 'there is no algorithm for truth' is contestable and may be a correctness risk, but contestability of a premise is not circularity. The derivation, insofar as one exists, is self-contained with respect to its citations, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Consciousness requires self-monitoring and an internal, updatable model of the external environment.
- domain assumption There is no algorithm for truth, so judging an LLM output as a hallucination is a subjective value judgment.
- domain assumption Next-token prediction that is indifferent to the truth of its output is equivalent to Frankfurtian bullshit.
Cite this review
Pith. "Pith review of A Primer on Large Language Models and their Limitations." pith.science (2026). https://pith.science/paper/M432F3A7
@misc{pith2026241204503,
author = {Pith},
title = {Pith review of: A Primer on Large Language Models and their Limitations},
year = {2026},
howpublished = {\url{https://pith.science/paper/M432F3A7}},
note = {Machine review of arXiv:2412.04503}
}
read the original abstract
This paper provides a primer on Large Language Models (LLMs) and identifies their strengths, limitations, applications and research directions. It is intended to be useful to those in academia and industry who are interested in gaining an understanding of the key LLM concepts and technologies, and in utilising this knowledge in both day to day tasks and in more complex scenarios where this technology can enhance current practices and processes.
Figures
Figures from the paper (16 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
Kevin D Ashley (2017): Artificial intelligence and legal analytics: New tools for law practice in the digital age, 6th print. edition. Cambridge Univ Press, CAMBRIDGE
work page 2017
-
[3]
Technical Report, University of Oxford
Mohamed Baioumy & Alex Cheema (2024): AI x Crypto Primer. Technical Report, University of Oxford. Available at https://alexcheema.github.io/AIxCryptoPrimer.pdf
work page 2024
-
[4]
Available at https://arxiv.org/abs/2304.12210
Randall Balestriero, Mark Ibrahim, Vlad Sobal, Ari Morcos, Shashank Shekhar, Tom Goldstein, Florian Bordes, Adrien Bardes, Gregoire Mialon, Yuandong Tian, Avi Schwarzschild, Andrew Gor- don Wilson, Jonas Geiping, Quentin Garrido, Pierre Fernandez, Amir Bar, Hamed Pirsiavash, Yann LeCun & Micah Goldblum (2023): A Cookbook of Self-Supervised Learning . Avai...
arXiv 2023
-
[5]
Available at https://www.deeplearning.ai/courses/ generative-ai-with-llms/
Antje Barth, Chris Fregly, Shelbee Eigenbrode & Mike Chambers: Generative AI with LLMs - DeepLearning.AI . Available at https://www.deeplearning.ai/courses/ generative-ai-with-llms/
-
[6]
(2017): Tfx: A tensorflow-based production- scale machine learning platform
Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc et al. (2017): Tfx: A tensorflow-based production- scale machine learning platform. In: Proceedings of the 23rd ACM SIGKDD international confer- ence on knowledge discovery and data mining, pp. 1387–1395
work page 2017
-
[7]
Celeste Biever (2023): ChatGPT broke the Turing test — the race is on for new ways to assess AI . Nature 619(7971), pp. 686–689, doi:10.1038/d41586-023-02361-7. Available at https://www. nature.com/articles/d41586-023-02361-7
-
[8]
Seth Bloomberg (2024): Dissecting the Intersection of AI and Crypto . Technical Report, Messari. Available at https://messari.io/report/ dissecting-the-intersection-of-ai-and-crypto
work page 2024
Show all 78 references
-
[9]
Available at https://arxiv.org/abs/2005.14165
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhari- wal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, 11ht...
2020 arXiv
-
[10]
In Tung-Hung Su & Jia-Horng Kao, editors: Artificial Intelligence, Machine Learning, and Deep Learning in Precision Medicine in Liver Diseases , Academic Press, pp
Alicia Chu, Liza Rachel Mathews & Kun-Hsing Yu (2023): Chapter 1 - Artificial intelligence in health care: past and present. In Tung-Hung Su & Jia-Horng Kao, editors: Artificial Intelligence, Machine Learning, and Deep Learning in Precision Medicine in Liver Diseases , Academi...
2023 doi
-
[11]
Le & Christopher D
Kevin Clark, Minh-Thang Luong, Quoc V . Le & Christopher D. Manning (2020): ELECTRA: Pre- training Text Encoders as Discriminators Rather Than Generators. Available at https://arxiv. org/abs/2003.10555
2020 arXiv
-
[12]
arXiv:1810.04805
Jacob Devlin, Ming-Wei Chang, Kenton Lee & Kristina Toutanova (2019): BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805
2019 arXiv
-
[13]
Science 384(6702), pp
Tyna Eloundou, Sam Manning, Pamela Mishkin & Daniel Rock (2024): GPTs are GPTs: Labor market impact potential of LLMs. Science 384(6702), pp. 1306–1308, doi:10.1126/science.adj0998. Available at https://www.science.org/doi/10.1126/science.adj0998
2024 doi
-
[14]
TechCrunch
Darrell Etherington (2021): MIT researchers develop a new ‘liquid’ neural network that’s better at adapting to new info . TechCrunch. Available at https://techcrunch.com/2021/01/ 28/mit-researchers-develop-a-new-liquid-neural-network-thats-better-at-\ adapting-to-new-info
2021
-
[15]
Princeton University Press, Princeton
Harry G Frankfurt (2009): On Bullshit. Princeton University Press, Princeton
2009
-
[16]
Cureus 15(9), doi:https://doi.org/10.7759/cureus.45473
Z ´u˜niga Salazar Gabriel, Diego Z ´u˜niga, Carlos L Vindel, Ana M Yoong, Hincapie Sofia, Ana B Z´u˜niga, Paula Z ´u˜niga, Erin Salazar & Z ´u˜niga Byron (2023): Efficacy of AI Chats to Determine an Emergency: A Comparison Between OpenAI’s ChatGPT, Google Bard, and Microsoft B...
2023 doi
-
[17]
Computational Linguistics, pp
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Der- noncourt, Tong Yu, Ruiyi Zhang & Nesreen K Ahmed (2024): Bias and fairness in large language models: A survey. Computational Linguistics, pp. 1–79
2024
-
[18]
In: International conference on machine learning, PMLR, pp
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat & Mingwei Chang (2020): Retrieval aug- mented language model pre-training. In: International conference on machine learning, PMLR, pp. 3929–3938
2020
-
[19]
Hannigan, Ian P
Timothy R. Hannigan, Ian P. McCarthy & Andr ´e Spicer (2024): Beware of botshit: How to manage the epistemic risks of generative chatbots . Business Horizons 67(5), pp. 471–486, doi:https://doi.org/10.1016/j.bushor.2024.03.001. Available at https://www.sciencedirect. com/scien...
2024 doi
-
[20]
AI 4(2), p
Cheng Hao-Wen (2023): Challenges and Limitations of ChatGPT and Artificial Intelligence for Scientific Research: A Perspective from Organic Materials. AI 4(2), p. 401, doi:10.3390/ai4020021
2023 doi
-
[21]
Ethics and Information Technology 26(2), p
Michael Townsen Hicks, James Humphries & Joe Slater (2024): ChatGPT is bullshit. Ethics and Information Technology 26(2), p. 38, doi:10.1007/s10676-024-09775-5. Available at https:// link.springer.com/10.1007/s10676-024-09775-5
2024 doi
-
[22]
Available at https://arxiv.org/abs/1503.02531
Geoffrey Hinton, Oriol Vinyals & Jeff Dean (2015): Distilling the Knowledge in a Neural Network. Available at https://arxiv.org/abs/1503.02531. S. Johnson & D. Hyland-Wood 29
2015 arXiv
-
[23]
International journal of uncertainty, fuzziness, and knowledge-based sys- tems 6(2), pp
Sepp Hochreiter (1998): The Vanishing Gradient Problem During Learning Recurrent Neural Nets and Problem Solutions. International journal of uncertainty, fuzziness, and knowledge-based sys- tems 6(2), pp. 107–116, doi:10.1142/S0218488598000094
1998 doi
-
[24]
arXiv (Cornell University) , doi:10.48550/arxiv.2203.15556
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osin...
-
[25]
Pattichis & Douglas B
Andreas Holzinger, Chris Biemann, Constantinos S. Pattichis & Douglas B. Kell (2017): What do we need to build explainable AI systems for the medical domain? Available at https://arxiv. org/abs/1712.09923
2017 arXiv
-
[26]
arXiv.org
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, An- drea Gesmundo, Mona Attariyan & Sylvain Gelly (2019): Parameter-Efficient Transfer Learning for NLP. arXiv.org. Available at http://arxiv.org/abs/1902.00751
2019 arXiv
-
[27]
Available at http://arxiv.org/abs/2106.09685
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang & Weizhu Chen (2021): LoRA: Low-Rank Adaptation of Large Language Models , doi:10.48550/arXiv.2106.09685. Available at http://arxiv.org/abs/2106.09685
-
[28]
Jelinek (1976): Continuous speech recognition by statistical methods
F. Jelinek (1976): Continuous speech recognition by statistical methods. Proceedings of the IEEE 64(4), pp. 532–556, doi:10.1109/PROC.1976.10159. Available at https://ieeexplore.ieee. org/document/1454428. Conference Name: Proceedings of the IEEE
1976
-
[29]
Available at https://keras.io/about/
Keras Team: Keras documentation: About Keras 3. Available at https://keras.io/about/
-
[30]
Proceedings of the National Academy of Sciences 114(13), pp
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, An- drei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran & Raia Hadsell (2017): Overcoming catastrophic forgett...
2017 doi
- [31]
- [32]
-
[33]
Progress in Biophysics and Molecular Biology 190, pp
Robert Lawrence Kuhn (2024): A landscape of consciousness: Toward a taxonomy of ex- planations and implications . Progress in Biophysics and Molecular Biology 190, pp. 28– 169, doi:10.1016/j.pbiomolbio.2023.12.003. Available athttps://linkinghub.elsevier.com/ retrieve/pii/S007...
2024 doi
-
[34]
arXiv:1910.13461
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov & Luke Zettlemoyer (2019): BART: Denoising Sequence-to-Sequence Pre- training for Natural Language Generation, Translation, and Comprehension. arXiv:1910.13461
2019 arXiv
- [35]
-
[36]
Available at https://arxiv.org/abs/2101.00190
Xiang Lisa Li & Percy Liang (2021): Prefix-Tuning: Optimizing Continuous Prompts for Genera- tion. Available at https://arxiv.org/abs/2101.00190. 30 LLM Primer
2021 arXiv
-
[37]
Advances in neural information processing systems 31
Yuan Li, Xiaodan Liang, Zhiting Hu & Eric P Xing (2018): Hybrid retrieval-generation reinforced agent for medical image report generation. Advances in neural information processing systems 31
2018
- [38]
-
[39]
Available at https://arxiv.org/abs/ 2401.15670
Jianqiao Lu, Wanjun Zhong, Yufei Wang, Zhijiang Guo, Qi Zhu, Wenyong Huang, Yanlin Wang, Fei Mi, Baojun Wang, Yasheng Wang, Lifeng Shang, Xin Jiang & Qun Liu (2024):YODA: Teacher- Student Progressive Learning for Language Models . Available at https://arxiv.org/abs/ 2401.15670
2024 arXiv
-
[40]
Available at https://arxiv.org/abs/2308.08747
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou & Yue Zhang (2024): An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning. Available at https://arxiv.org/abs/2308.08747
2024 arXiv
-
[41]
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson & Blaise Ag ¨uera y Arcas (2023): Communication-Efficient Learning of Deep Networks from Decentralized Data
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson & Blaise Ag ¨uera y Arcas (2023): Communication-Efficient Learning of Deep Networks from Decentralized Data . Available at https://arxiv.org/abs/1602.05629
2023 arXiv
-
[42]
Jackson (2024): A Turing test of whether AI chatbots are behaviorally similar to humans
Qiaozhu Mei, Yutong Xie, Walter Yuan & Matthew O. Jackson (2024): A Turing test of whether AI chatbots are behaviorally similar to humans . Proceedings of the National Academy of Sci- ences 121(9), doi:10.1073/pnas.2313925121. Available at https://www.pnas.org/doi/10. 1073/pna...
2024 doi
-
[43]
In: Image Analysis and Processing, ICIAP 2022 Workshops, PT II, Lecture Notes in Computer Science13374, Springer Nature, CHAM, pp
Gabriele Merlin, Vincenzo Lomonaco, Andrea Cossu, Antonio Carta & Davide Bacciu (2022): Practical Recommendations for Replay-Based Continual Learning Methods . In: Image Analysis and Processing, ICIAP 2022 Workshops, PT II, Lecture Notes in Computer Science13374, Springer Natu...
2022
-
[44]
Transactions of the International Society for Music Information Retrieval4, pp
Gianluca Micchi, Louis Bigo, Mathieu Giraud, Richard Groult & Florence Lev ´e (2021): I Keep Counting: An Experiment in Human/AI Co-creative Songwriting. Transactions of the International Society for Music Information Retrieval4, pp. 263+. Available athttp://dx.doi.org/10.5334...
2021
-
[45]
Available at https://arxiv.org/abs/ 2405.12630
Nicolo Micheletti, Samuel Belkadi, Lifeng Han & Goran Nenadic (2024): Exploration of Masked and Causal Language Modelling for Text Generation . Available at https://arxiv.org/abs/ 2405.12630
2024 arXiv
-
[46]
In: 2023 3rd International Conference on Innovative Mechanisms for Industry Applications (ICIMIA), IEEE, pp
G Mohan, G Satish, Harshal Patil, Vipul Vekariya, L Natrayan & Amit Barve (2023): AI-Powered Chatbot for Bridging Language Barriers with Translation . In: 2023 3rd International Conference on Innovative Mechanisms for Industry Applications (ICIMIA), IEEE, pp. 1559–1565
2023
-
[47]
In: Proceedings of the 38th ACM International Conference on Supercomputing , ICS ’24, Association for Computing Machinery, New York, NY , USA, pp
Baorun Mu, Christina Giannoula, Shang Wang & Gennady Pekhimenko (2024): Sylva: Sparse Embedded Adapters via Hierarchical Approximate Second-Order Information. In: Proceedings of the 38th ACM International Conference on Supercomputing , ICS ’24, Association for Computing Machin...
2024
-
[48]
Nikita Nangia, Clara Vania, Rasika Bhalerao & Samuel R Bowman (2020): CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models . In Bonnie Web- ber, Trevor Cohn, Yulan He & Yang Liu, editors: Proceedings of the 2020 Conference on Em- pirical Metho...
2020 doi
-
[49]
Available at https://arxiv.org/abs/2307.06435
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes & Ajmal Mian (2024):A Comprehensive Overview of Large Language Models. Available at https://arxiv.org/abs/2307.06435
2024 arXiv
-
[50]
Available at https://www.deeplearning.ai/
Andrew Ng: DeepLearning.AI. Available at https://www.deeplearning.ai/
-
[51]
Available at https://arxiv.org/ abs/2408.01505
Lin Ning, Harsh Lara, Meiqi Guo & Abhinav Rastogi (2024): MoDE: Effective Multi-task Param- eter Efficient Fine-Tuning with a Mixture of Dyadic Experts . Available at https://arxiv.org/ abs/2408.01505
2024 arXiv
-
[52]
Available at https://arxiv.org/abs/2203.02155
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike & Rya...
2022 arXiv
-
[53]
Computers and Electronics in Agri- culture 225, p
Hercules Panoutsopoulos, Borja Espejo-Garcia, Stephan Raaijmakers, Xu Wang, Spyros Foun- tas & Christopher Brewster (2024): Investigating the effect of different fine-tuning configura- tion scenarios on agricultural term extraction using BERT . Computers and Electronics in Agr...
2024
-
[54]
Available at https://arxiv.org/abs/2408.13296
Venkatesh Balavadhani Parthasarathy, Ahtsham Zafar, Aafaq Khan & Arsalan Shahid (2024): The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Tech- nologies, Research, Best Practices, Applied Research Challenges and Opportunities . Availa...
2024 arXiv
-
[55]
In: EACL 2021 - 16th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, pp
Jonas Pfeiffer, Aishwarya Kamath, Andreas R ¨uckl´e, Kyunghyun Cho & Iryna Gurevych (2021): AdapterFusion: Non-destructive task composition for transfer learning . In: EACL 2021 - 16th Conference of the European Chapter of the Association for Computational Linguistics, Proceed...
2021
-
[56]
Available at http://arxiv.org/abs/1910
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li & Peter J Liu (2023):Exploring the Limits of Transfer Learning with a Unified Text-to- Text Transformer, doi:10.48550/arXiv.1910.10683. Available at http://arxiv.org/abs/...
- [57]
-
[58]
Progress in Artificial Intelligence 9(4), pp
Mat ´ıas Roodschild, Jorge Gotay Sardi ˜nas & Adri ´an Will (2020): A new approach for the vanishing gradient problem on sigmoid activation . Progress in Artificial Intelligence 9(4), pp. 351–360, doi:10.1007/s13748-020-00218-y. Available at https://doi.org/10.1007/ s13748-020-00218-y
2020 doi
-
[59]
Available at http://arxiv.org/abs/1606.04671
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu & Raia Hadsell (2022): Progressive Neural Networks , doi:10.48550/arXiv.1606.04671. Available at http://arxiv.org/abs/1606.04671
-
[60]
In: 2024 IEEE Conference on Artificial In- telligence (CAI), IEEE, Singapore, Singapore, pp
Nicolai Schoch & Mario Hoernicke (2024): NL2IBE – Ontology-controlled Transformation of Nat- ural Language into Formalized Engineering Artefacts. In: 2024 IEEE Conference on Artificial In- telligence (CAI), IEEE, Singapore, Singapore, pp. 997–1004, doi:10.1109/CAI59869.2024.00...
2024
-
[61]
Nature Communications 15(1), p
Palistha Shrestha, Jeevan Kandel, Hilal Tayara & Kil To Chong (2024):Post-translational modifica- tion prediction via prompt-based fine-tuning of a GPT-2 model. Nature Communications 15(1), p. 32 LLM Primer 6699, doi:10.1038/s41467-024-51071-9. Available at https://www.proques...
2024 doi
-
[62]
Available at http://arxiv.org/abs/2305.17493
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot & Ross Anderson (2024): The Curse of Recursion: Training on Generated Data Makes Models Forget . Available at http://arxiv.org/abs/2305.17493. ArXiv:2305.17493 [cs]
2024 arXiv
-
[63]
Pearson, Hoboken, New Jersey
Ross Smith, Mayte Cubino Gonzalez & Emily McKeon (2024): The AI Revolution in Customer Ser- vice and Support: A Practical Guide to Impactful Deployment of AI to Best Serve Your Customers, [first edi edition. Pearson, Hoboken, New Jersey
2024
-
[64]
James W. A. Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, Michael S. A. Graziano & Cristina Becchio (2024): Testing theory of mind in large language models and humans . N...
2024 doi
- [65]
-
[66]
Available at https://arxiv.org/abs/2302.13971
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave & Guillaume Lample (2023): LLaMA: Open and Efficient Foundation L...
2023 arXiv
-
[67]
Gomez, Lukasz Kaiser & Illia Polosukhin (2023): Attention Is All You Need , doi:10.48550/arXiv.1706.03762
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser & Illia Polosukhin (2023): Attention Is All You Need , doi:10.48550/arXiv.1706.03762. Available at http://arxiv.org/abs/1706.03762. ArXiv:1706.03762 [cs]
-
[68]
Available at http://arxiv.org/abs/2204.05832
Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay & Colin Raffel (2022): What Language Model Architecture and Pretraining Ob- jective Work Best for Zero-Shot Generalization? , doi:10.48550/arXiv.2204.05832. Available at http:/...
-
[69]
Avail- able at https://arxiv.org/abs/2109.01652
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai & Quoc V Le (2022): Finetuned Language Models Are Zero-Shot Learners. Avail- able at https://arxiv.org/abs/2109.01652
2022 arXiv
-
[70]
Available at https://arxiv.org/abs/2201.11903
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le & Denny Zhou (2023):Chain-of-Thought Prompting Elicits Reasoning in Large Language Mod- els. Available at https://arxiv.org/abs/2201.11903
2023 arXiv
-
[71]
Innovation in language learning and teaching, pp
David James Woo, Kai Guo & Sdenka Zobeida Salas-Pilco (2024): Writing creative stories with AI: learning designs for secondary school students . Innovation in language learning and teaching, pp. 1–13
2024
-
[72]
Available at https://arxiv.org/abs/2303.17564v3
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prab- hanjan Kambadur, David Rosenberg & Gideon Mann (2023): BloombergGPT: A Large Language Model for Finance. Available at https://arxiv.org/abs/2303.17564v3
2023 arXiv
-
[73]
Na- ture Machine Intelligence 5(12), pp
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie & Fangzhao Wu (2023): Defending ChatGPT against jailbreak attack via self-reminders . Na- ture Machine Intelligence 5(12), pp. 1486–1496, doi:10.1038/s42256-023-00765-8. Available at https://w...
2023 doi
-
[74]
Available at https://arxiv.org/abs/2010.11934
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua & Colin Raffel (2021): mT5: A massively multilingual pre-trained text-to-text transformer . Available at https://arxiv.org/abs/2010.11934
2021 arXiv
-
[75]
In: 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Comput- ing - NeurIPS Edition (EMC2-NIPS), IEEE, pp
Ofir Zafrir, Guy Boudoukh, Peter Izsak & Moshe Wasserblat (2019): Q8BERT: Quantized 8Bit BERT. In: 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Comput- ing - NeurIPS Edition (EMC2-NIPS), IEEE, pp. 36–39, doi:10.1109/emc2-nips53020.2019.00016. Availabl...
2019
-
[76]
In: Conference on Parsimony and Learning , PMLR, pp
Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee & Yi Ma (2024):Inves- tigating the catastrophic forgetting in multimodal large language model fine-tuning. In: Conference on Parsimony and Learning , PMLR, pp. 202–227. Available at https://proceedings.mlr. ...
2024
-
[77]
Available at http://arxiv.org/abs/2303.18223
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie & Ji-Rong Wen ...
-
[78]
Technovation 124, p
Araz Zirar, Syed Imran Ali & Nazrul Islam (2023): Worker and workplace Artificial In- telligence (AI) coexistence: Emerging themes and research agenda . Technovation 124, p. 102747, doi:https://doi.org/10.1016/j.technovation.2023.102747. Available at https://www. sciencedirect...
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.