REVIEW 4 major objections 5 minor 1 cited by
Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A formal definition and three-part taxonomy organize the scattered field of LLM robustness.
desk verdict A useful organizational survey with a real repository, but the formal definition at Eq. (1) is ill-posed and should be fixed or dropped before this is referee-ready. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is Eq. (1), a min–max formulation that treats robustness as the worst-case loss over a set $\Delta$ of perturbation operations, with the original loss, the perturbed loss, and a distance term $d(\cdot,\cdot)$ (e.g., KL divergence) balanced by hyperparameters $\alpha$ and $\beta$. The formula is intended to cover three facets at once — performance on clean inputs, reliability on perturbed inputs, and consistency between the two outputs — and the paper's taxonomy in Fig. 2(b) is derived from it: adversarial robustness corresponds to perturbations $\epsilon$ such as $X'=X+\delta$, attack prompts, or long contexts; OOD robustness corresponds to distribution shifts with a bounded distance $\eta$ and an explicit "Not Known" refusal option; robustness evaluation supplies datasets, metrics, and benchmarks to measure the objective. This machinery carries the survey because every section assignment and every collected paper is justified by which part of Eq. (1) it addresses.
What would settle it
If one collected a random sample of robustness papers published in the past year and found a positively reviewed work that cannot be placed in any of the three branches, or that falls in two branches without double counting, the taxonomy's claim to be a comprehensive partition would fail. A simpler check: instantiate Eq. (1) literally, repairing the unbalanced parentheses and specifying the distance function $d(\cdot,\cdot)$, and show whether the 'reliability' term $L(\mathrm{LLM}(X'),Y')$ for $Y'\neq Y$ is even well-defined as a loss.
Extended reading notes
Core claim
The paper's central claim is that LLM robustness should be understood as a model's ability to maintain performance, consistency, and reliability across prompt variations, and that this ability can be captured by a single formal objective: $$\mathrm{Eval}(\$\theta$)=\arg\min_\$\theta$ \max_{\epsilon\in\$\Delta$} \big[ L(\mathrm{LLM}(X),Y) + \$\alpha$ L(\mathrm{LLM}(X'),Y') + \$\beta$\, d(L(\mathrm{LLM}(X))\parallel L(\mathrm{LLM}(X'))) \big],$$ where $X',Y'$ are perturbed data, $\Delta$ is a set of perturbation operations, and hyperparameters $\alpha,\beta$ trade off performance, consistency, and reliability. On the basis of this definition and a comparison with ML robustness, the paper proposes a three-part topology — adversarial robustness, OOD robustness, and robustness evaluation — and reviews representative works in each, alongside datasets, benchmarks, and future directions. The paper also asserts it is the first survey devoted specifically to LLM robustness rather than treating it as a subsection of a general LLM survey.
Load-bearing premise
The survey's organization rests on the assumption that robustness can be captured by a single min–max formula with two balancing weights, and that the three-way split of the literature follows naturally from that formula.
Editorial extensions
If this is right
- Researchers gain a shared vocabulary: 'noise prompt,' 'noise decoding,' 'OOD detection,' 'PEFT methods,' and 'hallucination' become named branches of a single robustness topology rather than separate subfields.
- The comparison with ML robustness (input, tuning, output, application) gives a checklist for where LLM robustness research is needed, such as prompt quality, parameter-efficient tuning, and knowledge updating.
- The formal definition implies that an LLM that is robust in the paper's sense must simultaneously be good on clean inputs, stable under perturbation, and consistent across paraphrases — three properties that existing benchmarks often measure separately.
- The human-in-the-loop framework (annotator, expert, red team, evaluator) provides a concrete process for continuously finding and patching robustness failures, and the future-directions table offers a chronological roadmap to 2029.
Reading between the lines
- Editorial inference: the taxonomy is a claim about how the literature clusters, not a theorem; a different perturbation typology (for example, one organized by attack surface rather than by input stage) would produce a different survey with equal plausibility.
- Editorial inference: Eq. (1) could be turned into a practical evaluation recipe, because fixing $\alpha$ and $\beta$ and sampling perturbations from $\Delta$ defines a family of robustness scores that the field does not yet standardize.
- Editorial inference: the same three-branch structure could be applied to multimodal and agentic LLMs, where the 'prompt' becomes a trajectory of observations and actions; whether the topology survives that extension is an open question.
- Editorial inference: the paper's own comparison with ML robustness implies that 'Not Known' refusal is a form of robustness, which points toward a testable extension — systems that learn to abstain under distribution shift may score higher on the paper's objective than systems that always answer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey of robustness in large language models. It claims to be the first comprehensive survey focused on LLM robustness, proposes a formal definition of LLM robustness in Eq. (1), and organizes the literature into three parts: adversarial robustness, OOD robustness, and robustness evaluation. It also provides a companion GitHub repository and discusses future directions including human-in-the-loop evaluation.
Significance. If the survey's indexing and taxonomy prove reliable, it would give the community a useful entry point: the three-part division by perturbation type is a reasonable organization, and the companion repository is a practical asset. The discussion of human roles and causal-inference connections is a forward-looking addition. However, the formal foundation is currently not rigorous, and several cited summaries contain errors; because the paper's value as a survey depends on accurate indexing, these issues need to be fixed.
major comments (4)
- [Section 1, Eq. (1)] The equation cannot serve as a formal definition of LLM robustness. The expression Eval(theta) is an argmin-max training objective, not a property or metric that characterizes whether a given model is robust. The perturbation set Delta is never linked to a distribution over (X',Y'), so the maximization is not well defined. The distance d(.,.) is left unspecified, the KL term contains unbalanced parentheses and is written with one argument, and the 'reliability' condition (Y' != Y) does not define a loss. Because the paper claims that the three-branch taxonomy is based on this definition, this is a load-bearing issue. The authors should either replace Eq. (1) with a well-typed definition of robustness as a model property, or explicitly present the taxonomy as an organizational choice based on perturbation types, as the abstract already does.
- [Section 1, contributions bullet] The claim 'we are the first to concentrate on the LLM Robustness' is contradicted by the paper's own references. Reference [179] (Wang et al., 'On the robustness of ChatGPT: An adversarial and out-of-distribution perspective') is a robustness-focused study of LLMs, and reference [213] (Yuan et al., 'Revisiting out-of-distribution robustness in NLP: Benchmarks, analysis, and LLMs evaluations') explicitly surveys OOD robustness in LLMs. The novelty claim should be qualified or removed, and the related-work comparison should engage with these existing surveys directly.
- [Section 5.2.2, reference [37]] The text states that 'Esiobu et al. [37] proposes ROBUST, the first benchmark for evaluating open information extraction models in real-world scenarios,' but reference [37] is titled 'ROBBIE: Robust bias evaluation of large generative language models' and concerns bias evaluation, not open information extraction. The same benchmark is also listed in Table 6 as 'ROBUST [37]' under 'Open domain generalization,' which is inconsistent. This is a substantive misattribution in a survey whose contribution is reliable indexing.
- [Section 3.2 and Table 3, reference [167]] The text attributes ALiBi to 'Sun et al. [167]' and says 'Their proposed ALiBi [167]' has been shown to outperform other position embedding methods. Reference [167] is the paper 'A Length-Extrapolatable Transformer' by Sun et al.; ALiBi is by Press et al. [137]. This type of attribution error is material in a survey-as-index. The authors should systematically audit all inline citations and tables against the reference list.
minor comments (5)
- [References [104] and [105]] References [104] and [105] are the same paper, 'Lost in the Middle: How Language Models Use Long Contexts,' cited with different volume numbers for the same TACL article; they should be merged, and Section 3.2 should cite a single entry.
- [Throughout] There are numerous typos and grammatical errors, including 'wild-range' for 'wide-range,' 'unexpeted' for 'unexpected,' 'perturbated' for 'perturbed,' 'disturber' for 'disturb,' 'finishi' and 'usenormous' in Section 1, and 'fla' for 'flag' in Section 2.2.2. A careful proofreading pass is needed.
- [Section 2.2.1, Eq. (2)] Equation (2) writes epsilon in {X' = X + delta, X' = Attack(X), LongContext}, mixing a scalar perturbation symbol with input-transformation conditions; the set-membership notation should be clarified.
- [Section 4.4] The sentence 'there also exist other related work [34, 47, 47, 68, 98, 150, 171]' contains a duplicated reference [47] and should be de-duplicated.
- [Section 5.2.2] The phrase 'he refined robustness metrics' has an unclear antecedent and appears to be a pronoun error; it should be 'they' or a specific author name should be given.
Circularity Check
No significant circularity: the survey's definition and taxonomy are stipulated, not derived from the paper's own outputs.
full rationale
This is a survey paper, and its central contributions are a formal definition of LLM robustness (Eq. 1), a three-part taxonomy (Adversarial, OOD, Evaluation), a paper-collection protocol, and a companion GitHub repository. None of these involve fitting a parameter and then predicting a closely related quantity, nor does any claimed result reduce by construction to an input. The formal definition in Eq. (1) is asserted rather than derived, and the taxonomy in Section 2.2 is explicitly justified by 'the comparison between ML robustness and LLM robustness' plus prior surveys [18, 160, 226] and the keyword statistics in Fig. 2(a), not by an inference from Eq. (1). The paper's claim to be 'the first to concentrate on the LLM Robustness' is a novelty assertion, not a load-bearing epistemic derivation, and the companion repository is an organizational asset rather than a circular support. The weaknesses noted by the reader, such as the unbalanced parentheses in Eq. (1), the unspecified distance d(·,·), and the loose mapping from the three bullets to the taxonomy, are correctness and rigor concerns about how well the definition is formalized; they are not instances of a result being equivalent to its own inputs. There is no self-citation chain invoked to forbid alternatives and no renamed empirical pattern presented as a derivation. Accordingly, the circularity score is 0.
Assumptions & free parameters
free parameters (2)
- alpha (Eq. 1)
- beta (Eq. 1)
assumptions (3)
- ad hoc to paper The min-max objective in Eq. (1) defines LLM robustness.
- domain assumption LLM robustness can be partitioned into adversarial robustness, OOD robustness, and evaluation.
- domain assumption The keyword-based search over selected venues yields a representative picture of the field.
Cite this review
Pith. "Pith review of Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions." pith.science (2026). https://pith.science/paper/QTDSRWLC
@misc{pith2026250611111,
author = {Pith},
title = {Pith review of: Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions},
year = {2026},
howpublished = {\url{https://pith.science/paper/QTDSRWLC}},
note = {Machine review of arXiv:2506.11111}
}
read the original abstract
Large Language Models (LLMs) have gained enormous attention in recent years due to their capability of understanding and generating natural languages. With the rapid development and wild-range applications (e.g., Agents, Embodied Intelligence), the robustness of LLMs has received increased attention. As the core brain of many AI applications, the robustness of LLMs requires that models should not only generate consistent contents, but also ensure the correctness and stability of generated content when dealing with unexpeted application scenarios (e.g., toxic prompts, limited noise domain data, outof-distribution (OOD) applications, etc). In this survey paper, we conduct a thorough review of the robustness of LLMs, aiming to provide a comprehensive terminology of concepts and methods around this field and facilitate the community. Specifically, we first give a formal definition of LLM robustness and present the collection protocol of this survey paper. Then, based on the types of perturbated inputs, we organize this survey from the following perspectives: 1) Adversarial Robustness: tackling the problem that prompts are manipulated intentionally, such as noise prompts, long context, data attack, etc; 2) OOD Robustness: dealing with the unexpected real-world application scenarios, such as OOD detection, zero-shot transferring, hallucinations, etc; 3) Evaluation of Robustness: summarizing the new evaluation datasets, metrics, and tools for verifying the robustness of LLMs. After reviewing the representative work from each perspective, we discuss and highlight future opportunities and research directions in this field. Meanwhile, we also organize related works and provide an easy-to-search project (https://github.com/zhangkunzk/Awesome-LLM-Robustness-papers) to support the community.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation
CoT probe-time gains arise primarily from lexical activation and short-range token co-occurrence rather than sentence-level logical derivation.
Reference graph
Works this paper leans on
-
[22]
Xuanting Chen, Junjie Ye, Can Zu, Nuo Xu, Rui Zheng, Minlong Peng, Jie Zhou, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023. How robust is gpt-3.5 to predecessors? a comprehensive study on language understanding tasks. arXiv preprint arXiv:2303.00293 (2023)
arXiv 2023
-
[41]
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024. Bias and fairness in large language models: A survey. Computational Linguistics (2024), 1–79
2024
-
[213]
Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji, Zhiyuan Liu, and Maosong Sun. 2023. Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and LLMs evaluations.Advances in Neural Information Processing Systems 36 (2023), 58478–58507
work page 2023
-
[104]
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 11 (2024), 157–173
2024
-
[105]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 12 (2024), 157–173
2024
-
[179]
Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, et al
-
[37]
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Smith. 2023. ROBBIE: Robust bias evaluation of large generative language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 3764–3814
2023
-
[167]
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei. 2023. A Length-Extrapolatable Transformer. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 14590–14604
2023
-
[137]
Ofir Press, Noah Smith, and Mike Lewis. 2022. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. In International Conference on Learning Representations
2022
Show all 246 references
-
[1]
Jameel Abdul Samadh, Mohammad Hanan Gani, Noor Hussein, Muhammad Uzair Khattak, Muhammad Muzammal Naseer, Fahad Shahbaz Khan, and Salman H Khan. 2024. Align your prompts: Test-time prompting with distribution alignment for zero-shot generalization. Advances in Neural Informati...
2024
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
2023 arXiv
-
[3]
Sravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, and R Venkatesh Babu. 2024. Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 23922–23932
2024
-
[4]
Dyah Adila, Changho Shin, Linrong Cai, and Frederic Sala. 2024. Zero-Shot Robustification of Zero-Shot Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=fCeUoDr9Tq
2024
-
[5]
Abhishek Aich, Calvin-Khang Ta, Akash Gupta, Chengyu Song, Srikanth Krishnamurthy, Salman Asif, and Amit Roy-Chowdhury. 2022. Gama: Generative adversarial multi-object scene attacks. Advances in Neural Information Processing Systems 35 (2022), 36914–36930
2022
-
[6]
Shengnan An, Zeqi Lin, Qiang Fu, Bei Chen, Nanning Zheng, Jian-Guang Lou, and Dongmei Zhang. 2023. How Do In-Context Examples Affect Compositional Generalization?. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers...
2023
-
[7]
Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, and Jian-Guang Lou. 2024. Make Your LLM Fully Utilize the Context. arXiv preprint arXiv:2404.16811 (2024)
2024 arXiv
-
[8]
Daman Arora, Himanshu Singh, et al. 2023. Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 7527–7543
2023
-
[9]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations
2023
-
[10]
Katherine Atwell, Mert Inan, Anthony B Sicilia, and Malihe Alikhani. 2024. Combining Discourse Coherence with Large Language Models for More Inclusive, Equitable, and Robust Task-Oriented Dialogue. In Proceedings of the 2024 Joint International Conference on Computational Ling...
2024
-
[11]
Reza Averly and Wei-Lun Chao. 2023. Unified out-of-distribution detection: A model-specific perspective. InProceedings of the IEEE/CVF International Conference on Computer Vision . 1453–1463
2023
-
[12]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073 (2022)
2022 arXiv
-
[13]
Sourya Basu, Pulkit Katdare, Prasanna Sattigeri, Vijil Chenthamarakshan, Katherine Driggs-Campbell, Payel Das, and Lav R Varshney
-
[14]
Francisco Bellas, Sara Guerreiro-Santalla, Martin Naya, and Richard J Duro. 2023. AI curriculum for European high schools: An embedded intelligence approach. International Journal of Artificial Intelligence in Education 33, 2 (2023), 399–426
2023
-
[15]
Adam Bouyamourn. 2023. Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 3181–3193
2023
-
[16]
Shreyas Bhat Brahmavar, Ashwin Srinivasan, Tirtharaj Dash, Sowmya Ramaswamy Krishnan, Lovekesh Vig, Arijit Roy, and Raviprasad Aduri. 2024. Generating Novel Leads for Drug Discovery using LLMs with Logical Feedback. In Proceedings of the AAAI Conference on Artificial Intellige...
2024
-
[17]
Houssem Ben Braiek and Foutse Khomh. 2024. Machine Learning Robustness: A Primer. arXiv preprint arXiv:2404.00897 (2024)
2024 arXiv
-
[18]
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45
2024
-
[19]
Huimin Chen, Chengyu Wang, Yanhao Wang, Cen Chen, and Yinggui Wang. 2024. TaiChi: Improving the Robustness of NLP Models by Seeking Common Ground While Reserving Differences. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resou...
2024
-
[20]
Hao Chen, Jindong Wang, Ankit Shah, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, and Bhiksha Raj. 2024. Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks. In The Twelfth International Conference on Learning Representations . https://openrevi...
2024
-
[21]
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. [n. d.]. Benchmarking large language models in retrieval-augmented generation. arXiv preprint arXiv:2309.01431 ([n. d.]). Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Article 111. Publication date: August 2024. 111:24 •...
2024 arXiv
-
[23]
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2024. LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models. In The Twelfth International Conference on Learning Representations
2024
-
[24]
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran. 2024. DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition ...
2024
-
[25]
Zining Chen, Weiqiu Wang, Zhicheng Zhao, Fei Su, Aidong Men, and Hongying Meng. 2024. PracticalDG: Perturbation Distillation on Vision-Language Models for Hybrid Domain Generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)....
2024
-
[26]
Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting Language Models to Compress Contexts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 3829–3846
2023
-
[27]
Zhixiang Chi, Li Gu, Tao Zhong, Huan Liu, YUANHAO YU, Konstantinos N Plataniotis, and Yang Wang. 2024. Adapting to Distribution Shift by Visual Domain Prompt Generation. In The Twelfth International Conference on Learning Representations . https://openreview. net/forum?id=sSaN4gxuEf
2024
-
[28]
Glass, and Pengcheng He
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2024. DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models. In The Twelfth International Conference on Learning Representations . https: //openreview.net/forum?id...
2024
-
[29]
Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. arXiv preprint arXiv:1809.05053 (2018)
2018 arXiv
-
[30]
Siddhartha Datta. 2022. Learn2weight: Parameter adaptation against similar-domain adversarial attacks. arXiv preprint arXiv:2205.07315 (2022)
2022 arXiv
-
[31]
Hillary Dawkins, Isar Nejadgholi, Daniel Gillis, and Judi McCuaig. 2024. Projective Methods for Mitigating Gender Bias in Pre-trained Language Models. arXiv preprint arXiv:2403.18803 (2024)
2024 arXiv
-
[32]
Nicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton, Elsbeth Turcan, and Kathleen Mckeown. 2023. Evaluation of African American Language Bias in Natural Language Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 6805–6824
2023
-
[33]
Zican Dong, Tianyi Tang, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Bamboo: A comprehensive benchmark for evaluating long text modeling capacities of large language models. arXiv preprint arXiv:2309.13345 (2023)
2023 arXiv
-
[34]
Yingpeng Du, Di Luo, Rui Yan, Xiaopei Wang, Hongzhi Liu, Hengshu Zhu, Yang Song, and Jie Zhang. 2024. Enhancing job recom- mendation through llm-based generative adversarial networks. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 8363–8371
2024
-
[35]
Haonan Duan, Adam Dziedzic, Nicolas Papernot, and Franziska Boenisch. 2023. Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models. InThirty-seventh Conference on Neural Information Processing Systems. https://openreview. net/forum?id=u6Xv3FuF8N
2023
-
[36]
Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan. 2022. A survey of embodied ai: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence 6, 2 (2022), 230–244
2022
-
[38]
Yu Fei, Yifan Hou, Zeming Chen, and Antoine Bosselut. 2023. Mitigating Label Biases for In-context Learning. InProceedings Of The 61St Annual Meeting Of The Association For Computational Linguistics (Acl 2023): Long Papers, Vol 1 . Assoc Computational Linguistics-Acl, 14014–14031
2023
-
[39]
Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. 2023. WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long ...
2023
-
[40]
Steinunn Rut Friðriksdóttir and Hafsteinn Einarsson. 2024. Gendered Grammar or Ingrained Bias? Exploring Gender Bias in Icelandic Language Models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-CO...
2024
-
[42]
Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024. ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving. InThe Twelfth International Conference on Learning Representations. 1–34. Proc. ACM Meas. Ana...
2024
-
[43]
Shreya Goyal, Sumanth Doddapaneni, Mitesh M Khapra, and Balaraman Ravindran. 2023. A survey of adversarial defenses and robustness in nlp. Comput. Surveys 55, 14s (2023), 1–39
2023
-
[44]
Sachin Goyal, Ananya Kumar, Sankalp Garg, Zico Kolter, and Aditi Raghunathan. 2023. Finetune like you pretrain: Improved finetuning of zero-shot vision models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 19338–19347
2023
-
[45]
Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 (2023)
2023 arXiv
-
[46]
Xinyan Guan, Yanjiang Liu, Hongyu Lin, Yaojie Lu, Ben He, Xianpei Han, and Le Sun. 2024. Mitigating large language model hallucinations via autonomous knowledge graph-based retrofitting. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 18126–18134
2024
-
[47]
Anisha Gunjal, Jihan Yin, and Erhan Bas. 2024. Detecting and preventing hallucinations in large vision language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 18135–18143
2024
-
[48]
Prakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri, Maxine Eskenazi, and Jeffrey P Bigham. 2022. InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pr...
2022
-
[49]
Fifty Shades of Bias
Rishav Hada, Agrima Seth, Harshita Diddee, and Kalika Bali. 2023. “Fifty Shades of Bias”: Normative Ratings of Gender Bias in GPT Generated English Text. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 1862–1876
2023
-
[50]
Dongchen Han, Xuran Pan, Yizeng Han, Shiji Song, and Gao Huang. 2023. Flatten transformer: Vision transformer using focused linear attention. In Proceedings of the IEEE/CVF international conference on computer vision . 5961–5971
2023
-
[51]
Jinwei Han, Zhiwen Lin, Zhongyisun Sun, Yingguo Gao, Ke Yan, Shouhong Ding, Yuan Gao, and Gui-Song Xia. 2024. Anchor-based Robust Finetuning of Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 26919–26928
2024
-
[52]
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu. 2023. Reasoning with Language Model is Planning with World Model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 8154–8173
2023
-
[53]
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2021. Self-attention attribution: Interpreting information interactions inside transformer. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 12963–12971
2021
-
[54]
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In Proceedings of the 60th Annual Meeting of the Association for Computational...
2022
-
[55]
Zexue He, Yu Wang, An Yan, Yao Liu, Eric Chang, Amilcare Gentili, Julian McAuley, and Chun-nan Hsu. 2023. MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model Evaluation. In Proceedings of the 2023 Conference on Empirical Methods in Natural...
2023
-
[56]
Peter Henderson, Eric Mitchell, Christopher Manning, Dan Jurafsky, and Chelsea Finn. 2023. Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society . 287–296
2023
-
[57]
Eugenio Herrera-Berg, Tomás Browne, Pablo León-Villagrá, Marc-Lluís Vives, and Cristian Calderon. 2023. Large Language Models are biased to overestimate profoundness. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 9653–9661
2023
-
[58]
Cheng-Yu Hsieh, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long T Le, Abhishek Kumar, James Glass, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, et al. 2024. Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization. arXiv preprint arXiv:...
2024 arXiv
-
[59]
Chengang Hu, Xiao Liu, and Yansong Feng. 2023. DiNeR: A Large Realistic Dataset for Evaluating Compositional Generalization. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 14938–14947
2023
-
[60]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021)
2021 arXiv
-
[61]
Zeyi Huang, Andy Zhou, Zijian Ling, Mu Cai, Haohan Wang, and Yong Jae Lee. 2023. A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language Guidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 11685–11695
2023
-
[62]
Shima Imani, Liang Du, and Harsh Shrivastava. 2023. Mathprompter: Mathematical reasoning using large language models. arXiv preprint arXiv:2303.05398 (2023)
2023 arXiv
-
[63]
Hamish Ivison, Akshita Bhagia, Yizhong Wang, Hannaneh Hajishirzi, and Matthew E Peters. 2023. HINT: Hypernetwork Instruction Tuning for Efficient Zero-and Few-Shot Generalisation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volum...
2023
-
[64]
Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based neural structured learning for sequential question answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1821–1831
2017
-
[65]
Akshita Jha, Aida Mostafazadeh Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prabhakaran, and Sunipa Dev. 2023. SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models. In Proceedings of the 61st Annual Meeting Proc. ACM Meas. Anal. Com...
2023
-
[66]
Akshita Jha and Chandan K Reddy. 2023. Codeattack: Code-based adversarial attacks for pre-trained programming language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 14892–14900
2023
-
[67]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088 (2024)
2024 arXiv
-
[68]
Chaoya Jiang, Haiyang Xu, Mengfan Dong, Jiaxing Chen, Wei Ye, Ming Yan, Qinghao Ye, Ji Zhang, Fei Huang, and Shikun Zhang. 2024. Hallucination Augmented Contrastive Learning for Multimodal Large Language Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...
2024
-
[69]
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023. Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression. arXiv preprint arXiv:2310.06839 (2023)
2023 arXiv
-
[70]
Shuoran Jiang, Qingcai Chen, Yang Xiang, Youcheng Pan, and Yukang Lin. 2024. Linguistic Rule Induction Improves Adversarial and OOD Robustness in Large Language Models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources a...
2024
-
[71]
Xue Jiang, Feng Liu, Zhen Fang, Hong Chen, Tongliang Liu, Feng Zheng, and Bo Han. 2024. Negative Label Guided OOD Detection with Pretrained Vision-Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/ forum?id=xUO1HXz4an
2024
-
[72]
Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is ChatGPT a good translator? A preliminary study. arXiv preprint arXiv:2301.08745 1, 10 (2023)
2023 arXiv
-
[73]
Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, LYU Zhiheng, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman- Weiner, Mrinmaya Sachan, et al. 2023. Cladder: Assessing causal reasoning in language models. In Thirty-seventh conference on neural information proc...
2023
-
[74]
Erik Jones, Hamid Palangi, Clarisse Simões Ribeiro, Varun Chandrasekaran, Subhabrata Mukherjee, Arindam Mitra, Ahmed Hassan Awadallah, and Ece Kamar. 2024. Teaching Language Models to Hallucinate Less with Synthetic Tasks. InThe Twelfth International Conference on Learning Rep...
2024
-
[75]
Adilbek Karmanov, Dayan Guan, Shijian Lu, Abdulmotaleb El Saddik, and Eric Xing. 2024. Efficient Test-Time Adaptation of Vision- Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14162–14171
2024
-
[76]
Aly Kassem, Omar Mahmoud, and Sherif Saad. 2023. Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 4360–4379
2023
-
[77]
Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. 2023. A survey of reinforcement learning from human feedback. arXiv preprint arXiv:2312.14925 (2023)
2023
-
[78]
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2023. ProPILE: Probing Privacy Leakage in Large Language Models. In Thirty-seventh Conference on Neural Information Processing Systems . https://openreview.net/forum?id= QkLpGxUboF
2023
-
[79]
Ching-Yun Ko, Pin-Yu Chen, Payel Das, Yung-Sung Chuang, and Luca Daniel. 2023. On Robustness-Accuracy Characterization of Large Language Models using Synthetic Datasets. (2023)
2023
-
[80]
Giorgi Kokaia, Pratyush Sinha, Yutong Jiang, and Nozha Boujemaa. 2023. Writing your own book: A method for going from closed to open book QA to improve robustness and performance of smaller LLMs. arXiv preprint arXiv:2305.11334 (2023)
2023 arXiv
-
[81]
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. In Proceedings of the 2016 conference of the north american chapter of the association for computational linguistics: human language technologies....
2016
-
[82]
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. 2023. Pretraining language models with human preferences. In International Conference on Machine Learning . PMLR, 17506–17533
2023
-
[83]
Mayank Kothyari, Dhruva Dhingra, Sunita Sarawagi, and Soumen Chakrabarti. 2023. CRUSH4SQL: Collective Retrieval Using Schema Hallucination For Text2SQL. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 14054–14066
2023
-
[84]
Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. 2022. Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution. In International Conference on Learning Representations . https://openreview.net/ forum?id=UYneFzXSJWh
2022
-
[85]
Po-Nien Kung, Fan Yin, Di Wu, Kai-Wei Chang, and Nanyun Peng. 2023. Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 1813–1829
2023
-
[86]
Yu-Ju Lan and Nian-Shing Chen. 2024. Teachers’ agency in the era of LLM and generative AI. Educational Technology & Society 27, 1 (2024), I–XVIII. Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Article 111. Publication date: August 2024. Evaluating and Improving Robustne...
2024
-
[87]
Yoonho Lee, Annie S Chen, Fahim Tajwar, Ananya Kumar, Huaxiu Yao, Percy Liang, and Chelsea Finn. 2023. Surgical Fine-Tuning Improves Adaptation to Distribution Shifts. In The Eleventh International Conference on Learning Representations . https://openreview. net/forum?id=APuPRxjHvZ
2023
-
[88]
Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. 2024. Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2024
-
[89]
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen Mckeown, and William Yang Wang. 2022. SafeText: A Benchmark for Exploring Physical Safety in Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pro...
2022
-
[90]
Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, and Najoung Kim. 2023. SLOG: A Structural Generalization Benchmark for Semantic Parsing. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 3213–3232
2023
-
[91]
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023. HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 6449–6464
2023
-
[92]
Juncheng Li, Minghe Gao, Longhui Wei, Siliang Tang, Wenqiao Zhang, Mengze Li, Wei Ji, Qi Tian, Tat-Seng Chua, and Yueting Zhuang. 2023. Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language Models. In Proceedings of the IEEE/CVF International Conference on ...
2023
-
[93]
Lin Li, Haoyan Guan, Jianing Qiu, and Michael Spratling. 2024. One prompt word is enough to boost adversarial robustness for pre-trained vision-language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 24408–24419
2024
-
[94]
Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli. 2024. Functional Interpolation for Relative Positions improves Long Context Transformers. In The Twelfth International Co...
2024
-
[95]
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2022. Large Language Models Can Be Strong Differentially Private Learners. In International Conference on Learning Representations . https://openreview.net/forum?id=bVuP3ltATMz
2022
-
[96]
Xiao Li, Wei Zhang, Yining Liu, Zhanhao Hu, Bo Zhang, and Xiaolin Hu. 2024. Language-Driven Anchors for Zero-Shot Adversarial Robustness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 24686–24695
2024
-
[97]
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. Contrastive Decoding: Open-ended Text Generation as Optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational L...
2023
-
[98]
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Evaluating Object Hallucination in Large Vision-Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 292–305
2023
-
[99]
Yufei Li, Zexin Li, Yingfan Gao, and Cong Liu. 2023. White-box multi-objective adversarial attack on dialogue generation. arXiv preprint arXiv:2305.03655 (2023)
2023 arXiv
-
[100]
Jian Liang, Ran He, and Tieniu Tan. 2024. A comprehensive survey on test-time adaptation under distribution shifts. International Journal of Computer Vision (2024), 1–34
2024
-
[101]
Victoria Lin, Eli Ben-Michael, and Louis-Philippe Morency. [n. d.]. Optimizing Language Models for Human Preferences is a Causal Inference Problem. In The 40th Conference on Uncertainty in Artificial Intelligence
-
[102]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. DeepSeek-V3 Technical Report. arXiv preprint arXiv:2412.19437 (2024)
2024 arXiv
-
[103]
Bo Liu, Li-Ming Zhan, Zexin Lu, Yujie Feng, Lei Xue, and Xiao-Ming Wu. 2024. How Good Are LLMs at Out-of-Distribution Detection?. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 8211–8222
2024
-
[106]
Yiran Liu, Xiaoang Xu, Zhiyi Hou, and Yang Yu. 2024. Causality Based Front-door Defense Against Backdoor Attack on Language Models. In Forty-first International Conference on Machine Learning . https://openreview.net/forum?id=dmHHVcHFdM
2024
-
[107]
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023. Trustworthy LLMs: A survey and guideline for evaluating large language models’ alignment. arXiv preprint arXiv:2308.05374 (2023)
2023 arXiv
-
[108]
Dong Lu, Zhiqiang Wang, Teng Wang, Weili Guan, Hongchang Gao, and Feng Zheng. 2023. Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 102–111. Proc...
2023
-
[109]
Enzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du, Tao Jiang, Chao Hong, Shaowei Liu, Weiran He, Enming Yuan, Yuzhi Wang, et al
-
[110]
Xinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov, and Min-Yen Kan. 2023. SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 7787–7813
2023
-
[111]
Fan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian, Jiashi Feng, and Yi Yang. 2024. VISTA-LLAMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 13151–13160
2024
-
[112]
Rishabh Maheshwary, Saket Maheshwary, and Vikram Pudi. 2021. A context aware approach for generating natural language attacks. In Proceedings of the AAAI conference on artificial intelligence , Vol. 35. 15839–15840
2021
-
[113]
Rishabh Maheshwary, Saket Maheshwary, and Vikram Pudi. 2021. Generating natural language attacks in a hard label black box setting. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 13525–13533
2021
-
[114]
Potsawee Manakul, Adian Liusie, and Mark Gales. 2023. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 9004–9017
2023
-
[115]
Yuren Mao, Yuhang Ge, Yijiang Fan, Wenyi Xu, Yu Mi, Zhonghao Hu, and Yunjun Gao. 2024. A Survey on LoRA of Large Language Models. arXiv preprint arXiv:2407.11046 (2024)
2024 arXiv
-
[116]
Weikang Meng, Yadan Luo, Xin Li, Dongmei Jiang, and Zheng Zhang. 2025. PolaFormer: Polarity-aware Linear Attention for Vision Transformers. In The Thirteenth International Conference on Learning Representations
2025
-
[117]
Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2021. A diverse corpus for evaluating and developing English math word problem solvers. arXiv preprint arXiv:2106.15772 (2021)
2021 arXiv
-
[118]
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey. arXiv preprint arXiv:2402.06196 (2024)
2024 arXiv
-
[119]
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-Task Generalization via Natural Language Crowdsourcing Instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 3470–3487
2022
-
[120]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, et al. 2023. Crosslingual Generalization through Multitask Finetuning. In Proceedings of the 61st Annual Meeting of th...
2023
-
[121]
Manish Nagireddy, Lamogha Chiazor, Moninder Singh, and Ioana Baldini. 2024. Socialstigmaqa: A benchmark to uncover stigma amplification in generative language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21454–21462
2024
-
[122]
Muzammal Naseer, Ahmad Mahmood, Salman Khan, and Fahad Khan. 2023. Boosting Adversarial Transferability using Dynamic Cues. In The Eleventh International Conference on Learning Representations . https://openreview.net/forum?id=SZynfVLGd5
2023
-
[123]
Allen Nie, Yuhui Zhang, Atharva Shailesh Amdekar, Chris Piech, Tatsunori B Hashimoto, and Tobias Gerstenberg. 2024. Moca: Measuring human-language model alignment on causal and moral judgment tasks. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[124]
Ardavan Salehi Nobandegani, Kevin da Silva Castanheira, Timothy O’Donnell, and Thomas R Shultz. 2019. On Robustness: An Undervalued Dimension of Human Rationality.. In CogSci. 3327
2019
-
[125]
Changdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim, Minchul Shin, Jong-June Jeon, and Kyungwoo Song. 2024. Geodesic multi-modal mixup for robust fine-tuning. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[126]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35...
2022
-
[127]
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: long papers) . 1946–1958
2017
-
[128]
Giuseppe Paolo, Jonas Gonzalez-Billandon, and Balázs Kégl. 2024. A call for embodied AI. arXiv preprint arXiv:2402.03824 (2024)
2024 arXiv
-
[129]
Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables. arXiv preprint arXiv:1508.00305 (2015)
2015 arXiv
-
[130]
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? arXiv preprint arXiv:2103.07191 (2021)
2021 arXiv
-
[131]
Arkil Patel, Satwik Bhattamishra, Siva Reddy, and Dzmitry Bahdanau. 2023. MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel Interpretations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Proce...
2023
-
[132]
was it “stated
Roma Patel and Ellie Pavlick. 2021. “was it “stated” or was it “claimed”?: How linguistic bias affects generative language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 10080–10095
2021
-
[133]
Kellin Pelrine, Anne Imouza, Camille Thibault, Meilina Reksoprodjo, Caleb Gupta, Joel Christoph, Jean-François Godbout, and Reihaneh Rabbany. 2023. Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4. In Proceedings of the 2023 Conference on Empi...
2023
-
[134]
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Leon Derczynski, Xingjian Du, Matteo Grella, Kranthi Gv, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartłomiej Koptyra, Hay...
2023
-
[135]
Luiza Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker. 2023. On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 7595–7609
2023
-
[136]
Bodhisattwa Prasad Majumder, Zexue He, and Julian McAuley. 2022. InterFair: Debiasing with Natural Language Feedback for Fair Interpretable Predictions. arXiv e-prints (2022), arXiv–2210
2022
-
[138]
Ji Qi, Chuchun Zhang, Xiaozhi Wang, Kaisheng Zeng, Jifan Yu, Jinxin Liu, Lei Hou, Juanzi Li, and Xu Bin. 2023. Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction. In Proceedings of the 2023 Conference on Empirical Methods in Natura...
2023
-
[139]
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. 2024. Visual adversarial examples jailbreak aligned large language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21527–21536
2024
-
[140]
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2024. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!. In The Twelfth International Conference on Learning Representations . https://openrev...
2024
-
[141]
Yifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Ponti, and Shay B Cohen. 2023. Detecting and Mitigating Hallucinations in Multilingual Summarisation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 8914–8932
2023
-
[142]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023. Direct preference optimization: your language model is secretly a reward model. In Proceedings of the 37th International Conference on Neural Information Processing Sys...
2023
-
[143]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[144]
Vivek Ramanujan, Thao Nguyen, Sewoong Oh, Ali Farhadi, and Ludwig Schmidt. 2023. On the Connection between Pre-training Data Diversity and Fine-tuning Robustness. In Thirty-seventh Conference on Neural Information Processing Systems . https://openreview.net/ forum?id=2SScUiWUbn
2023
-
[145]
Michael Ryan, Tarek Naous, and Wei Xu. 2023. Revisiting non-English Text Simplification: A Unified Multilingual Benchmark. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 4898–4927
2023
-
[146]
Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyo...
2022
-
[147]
Sebastin Santy, Jenny T Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. NLPositionality: Characterizing Design Biases of Datasets and Models. In The 61st Annual Meeting Of The Association For Computational Linguistics
2023
-
[148]
Soumya Sanyal, Zeyi Liao, and Xiang Ren. 2022. RobustLR: A diagnostic benchmark for evaluating logical robustness of deductive reasoners. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 9614–9631
2022
-
[149]
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations. htt...
2024
-
[150]
Ziyu Shang, Wenjun Ke, Nana Xiu, Peng Wang, Jiajun Liu, Yanhui Li, Zhizhao Luo, and Ke Ji. 2024. Ontofact: Unveiling fantastic fact-skeleton of llms via ontology-driven reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 18934–18943
2024
-
[151]
NAN SHAO, Zefan Cai, Hanwei xu, Chonghua Liao, Yanan Zheng, and Zhilin Yang. 2023. Compositional Task Representations for Large Language Models. InThe Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=6axIMJA7ME3 Proc. ACM Meas. Ana...
2023
-
[152]
Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. 2024. Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=plmBsXHxgR
2024
-
[153]
Ke Shen. 2024. The Generalization and Robustness of Transformer-Based Language Models on Commonsense Reasoning. InProceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 23419–23420
2024
-
[154]
Atsushi Shirafuji, Yutaka Watanobe, Takumi Ito, Makoto Morishita, Yuki Nakamura, Yusuke Oda, and Jun Suzuki. 2023. Exploring the robustness of large language models for solving programming problems. arXiv preprint arXiv:2306.14583 (2023)
2023 arXiv
-
[155]
Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, and Chaowei Xiao. 2022. Test-time prompt tuning for zero-shot generalization in vision-language models. Advances in Neural Information Processing Systems 35 (2022), 14274–14289
2022
-
[156]
Yang Shu, Xingzhuo Guo, Jialong Wu, Ximei Wang, Jianmin Wang, and Mingsheng Long. 2023. Clipood: Generalizing clip to out-of-distributions. In International Conference on Machine Learning . PMLR, 31716–31731
2023
-
[157]
Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng, Danqi Chen, and He He. 2023. Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations. Association for Computational Linguistics
2023
-
[158]
Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel. 2023. The curious case of hallucinatory (un) answerability: Finding truths in the hidden states of over-confident large language models. In Proceedings of the 2023 Conference on Empirical Methods in N...
2023
-
[159]
Joe Stacey, Yonatan Belinkov, and Marek Rei. 2022. Supervising model attention with human explanations for robust natural language inference. In Proceedings of the AAAI conference on artificial intelligence , Vol. 36. 11349–11357
2022
-
[160]
Michal Štefánik. 2022. Methods for Estimating and improving robustness of language models. arXiv preprint arXiv:2206.08446 (2022)
2022 arXiv
-
[161]
Asa Cooper Stickland, Sailik Sengupta, Jason Krone, Saab Mansour, and He He. 2022. Robustification of multilingual language models to real-world noise in crosslingual zero-shot settings with robust contrastive pretraining. arXiv preprint arXiv:2210.04782 (2022)
2022 arXiv
-
[162]
Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schölkopf, and Mrinmaya Sachan. 2022. A causal framework to quantify the robustness of mathematical reasoning with language models. arXiv preprint arXiv:2210.12023 (2022)
2022 arXiv
-
[163]
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing 568 (2024), 127063
2024
-
[164]
Hao Sun and Mihaela van der Schaar. 2024. Inverse-RLignment: Inverse Reinforcement Learning from Demonstrations for LLM Alignment. arXiv preprint arXiv:2405.15624 (2024)
2024 arXiv
-
[165]
Jiuding Sun, Chantal Shaib, and Byron C Wallace. 2024. Evaluating the Zero-shot Robustness of Instruction-tuned Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=g9diuvxN6D
2024
-
[166]
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023. Retentive network: A successor to transformer for large language models. arXiv preprint arXiv:2307.08621 (2023)
2023 arXiv
-
[168]
Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. 2024. Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem. arXiv preprint arXiv:2403.03558 (2024)
2024 arXiv
-
[169]
Song Tang, Wenxin Su, Mao Ye, and Xiatian Zhu. 2024. Source-Free Domain Adaptation with Frozen Multimodal Foundation Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 23711–23720
2024
-
[170]
Junjiao Tian, Zecheng He, Xiaoliang Dai, Chih-Yao Ma, Yen-Cheng Liu, and Zsolt Kira. 2023. Trainable projected gradient method for robust fine-tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7836–7845
2023
-
[171]
Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn. 2024. Fine-Tuning Language Models for Factuality. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=WPZ2yPag4K
2024
-
[172]
Andrea Tocchetti, Lorenzo Corti, Agathe Balayn, Mireia Yurrita, Philip Lippmann, Marco Brambilla, and Jie Yang. 2022. AI robustness: a human-centered perspective on technological challenges and opportunities. Comput. Surveys (2022)
2022
-
[173]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[174]
Thiagarajan
Puja Trivedi, Danai Koutra, and Jayaraman J. Thiagarajan. 2023. A Closer Look at Model Adaptation using Feature Distortion and Simplicity Bias. In The Eleventh International Conference on Learning Representations . https://openreview.net/forum?id=wkg_b4-IwTZ
2023
-
[175]
Faizad Ullah, Ali Faheem, Ubaid Azam, Muhammad Sohaib Ayub, Faisal Kamiran, and Asim Karim. 2024. Detecting Cybercrimes in Accordance with Pakistani Law: Dataset and Evaluation Using PLMs. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, ...
2024
-
[176]
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning. In International Conference on Machine Learning . PMLR, 35413–35425
2023
-
[177]
Guangya Wan, Yuqi Wu, Mengxuan Hu, Zhixuan Chu, and Sheng Li. 2024. Bridging causal discovery and large language models: A comprehensive survey of integrative approaches and future directions. arXiv preprint arXiv:2402.11068 (2024). Proc. ACM Meas. Anal. Comput. Syst., Vol. 37...
2024 arXiv
-
[178]
Boxin Wang, Wei Ping, Chaowei Xiao, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Bo Li, Anima Anandkumar, and Bryan Catanzaro. 2022. Exploring the limits of domain-adaptive training for detoxifying large-scale language models. Advances in Neural Information Processing Systems 3...
2022
-
[180]
Shiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, et al. 2022. Recode: Robustness evaluation of code generation models. arXiv preprint arXiv:2212.10264 (2022)
2022 arXiv
-
[181]
Sibo Wang, Jie Zhang, Zheng Yuan, and Shiguang Shan. 2024. Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 24502–24511
2024
-
[182]
Tianlu Wang, Rohit Sridhar, Diyi Yang, and Xuezhi Wang. 2021. Identifying and mitigating spurious correlations for improving robustness in nlp models. arXiv preprint arXiv:2110.07736 (2021)
2021 arXiv
-
[183]
Xuezhi Wang, Haohan Wang, and Diyi Yang. 2021. Measure and improve robustness in NLP models: A survey. arXiv preprint arXiv:2112.08313 (2021)
2021 arXiv
-
[184]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations. 1–24
2023
-
[185]
Yihan Wang, Si Si, Daliang Li, Michal Lukasik, Felix Yu, Cho-Jui Hsieh, Inderjit S Dhillon, and Sanjiv Kumar. 2024. Two-stage LLM Fine-tuning with Less Specialization and More Generalization. In The Twelfth International Conference on Learning Representations . https://openrev...
2024
-
[186]
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How Does LLM Safety Training Fail?. In Thirty-seventh Conference on Neural Information Processing Systems . https://openreview.net/forum?id=jA235JGM09
2023
-
[187]
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682 (2022)
2022 arXiv
-
[188]
Zhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu, Tom Goldstein, and Yu-Gang Jiang. 2022. Towards transferable adversarial attacks on vision transformers. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 2668–2676
2022
-
[189]
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang. 2023. Unveiling the Implicit Toxicity in Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 1322–1338
2023
-
[190]
Anpeng Wu, Kun Kuang, Minqin Zhu, Yingrong Wang, Yujia Zheng, Kairong Han, Baohong Li, Guangyi Chen, Fei Wu, and Kun Zhang
-
[191]
Cheng-En Wu, Yu Tian, Haichao Yu, Heng Wang, Pedro Morgado, Yu Hen Hu, and Linjie Yang. 2023. Why Is Prompt Tuning for Vision-Language Models Robust to Noisy Labels?. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 15488–15497
2023
-
[192]
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023. Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 3909–3925
2023
-
[193]
Yu Xia, Tong Yu, Zhankui He, Handong Zhao, Julian McAuley, and Shuai Li. 2024. Aligning as Debiasing: Causality-Aware Alignment via Reinforcement Learning with Interventional Feedback. In Proceedings of the 2024 Conference of the North American Chapter of the Association for C...
2024
-
[194]
CoRR (2024)
Causality for Large Language Models. CoRR (2024)
2024
-
[195]
Yao Xiao, Ziyi Tang, Pengxu Wei, Cong Liu, and Liang Lin. 2023. Masked images are counterfactual samples for robust fine-tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 20301–20310
2023
-
[196]
Zi Xiong, Lizhi Qing, Yangyang Kang, Jiawei Liu, Hongsong Li, Changlong Sun, Xiaozhong Liu, and Wei Lu. 2024. Enhance Robustness of Language Models Against Variation Attack through Graph Integration. arXiv preprint arXiv:2404.12014 (2024)
2024 arXiv
-
[197]
Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, and Dongkuan Xu. 2023. Rewoo: Decoupling reasoning from observations for efficient augmented language models. arXiv preprint arXiv:2305.18323 (2023)
2023 arXiv
-
[198]
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li. 2024. BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models. In The Twelfth International Conference on Learning Representations . https: //openreview.net/forum?id=c...
2024
-
[199]
Weijia Xu, Batool Haider, and Saab Mansour. 2020. End-to-end slot alignment and recognition for cross-lingual NLU. arXiv preprint arXiv:2004.14353 (2020)
2020 arXiv
-
[200]
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, and Qian Lou. 2023. TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models. In Thirty-seventh Conference on Neural Information Processing Systems . https://openreview. net/forum?id=ZejTutd...
2023
-
[201]
Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian. 2024. RLCD: Reinforcement learning from contrastive distillation for LM alignment. In The Twelfth International Conference on Learning Representations
2024
-
[202]
Silei Xu, Shicheng Liu, Theo Culhane, Elizaveta Pertseva, Meng-Hsi Wu, Sina Semnani, and Monica Lam. 2023. Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata. In Proceedings of the 2023 Conference on Empirical Methods ...
2023
-
[203]
Yu Yang, Besmira Nushi, Hamid Palangi, and Baharan Mirzasoleiman. 2023. Mitigating spurious correlations in multi-modal models during fine-tuning. In International Conference on Machine Learning . PMLR, 39365–39379
2023
-
[204]
Ziqing Yang, Xinlei He, Zheng Li, Michael Backes, Mathias Humbert, Pascal Berrang, and Yang Zhang. 2023. Data poisoning attacks against multimodal encoders. In International Conference on Machine Learning . PMLR, 39299–39313
2023
-
[205]
Zonghan Yang and Yang Liu. 2022. On Robust Prefix-Tuning for Text Classification. In International Conference on Learning Representa- tions. https://openreview.net/forum?id=eBCmOocUejf
2022
-
[206]
Yuting Yang, Pei Huang, Feifei Ma, Juan Cao, and Jintao Li. 2024. PAD: A Robustness Enhancement Ensemble Method via Promoting Attention Diversity. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-CO...
2024
-
[207]
Ziyi Yin, Muchao Ye, Tianrong Zhang, Jiaqi Wang, Han Liu, Jinghui Chen, Ting Wang, and Fenglong Ma. 2024. VQAttack: Transferable Adversarial Attacks on Visual Question Answering via Pre-trained Models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 6755–6763
2024
-
[208]
Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A Inan, Gautam Kamath, Janardhan Kulkarni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, Sergey Yekhanin, and Huishuai Zhang. 2022. Differentially Private Fine-tuning of Language Models. In International Conference on ...
2022
-
[209]
Jifan Yu, Xiaohan Zhang, Yifan Xu, Xuanyu Lei, Zijun Yao, Jing Zhang, Lei Hou, and Juanzi Li. 2024. A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation. arXiv preprint arXiv:2404.03491 (2024)
2024 arXiv
-
[210]
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems 36 (2023), 11809–11822
2023
-
[211]
Zichun Yu, Chenyan Xiong, Shi Yu, and Zhiyuan Liu. 2023. Augmentation-Adapted Retriever Improves Generalization of Language Models as Generic Plug-In. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2421–2436
2023
-
[212]
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, YX Wei, Lean Wang, Zhiping Xiao, et al. 2025. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. arXiv preprint arXiv:2502.11089 (2025)
2025 arXiv
-
[214]
Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, and Tat-Seng Chua. 2024. RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback. In Proceedings of the IEEE/CV...
2024
-
[215]
Zhao Yukun, Yan Lingyong, Sun Weiwei, Xing Guoliang, Wang Shuaiqiang, Meng Chong, Cheng Zhicong, Ren Zhaochun, and Yin Dawei. 2024. Improving the robustness of large language models via consistency alignment. arXiv preprint arXiv:2403.14221 (2024)
2024 arXiv
-
[216]
Maxime Zanella and Ismail Ben Ayed. 2024. On the Test-Time Zero-Shot Generalization of Vision-Language Models: Do We Really Need Prompt Learning?. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 23783–23793
2024
-
[217]
Susskind, and Chen Huang
Yuhang Zang, Hanlin Goh, Joshua M. Susskind, and Chen Huang. 2024. Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id= PKICZXVY9M
2024
-
[218]
Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi
-
[219]
In The Twelfth International Conference on Learning Representations
Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=gfFVATffPd
-
[220]
Tianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng, Yue Zhang, Zheng Zhang, Chenghu Zhou, Xinbing Wang, and Luoyi Fu. 2023. Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Proc...
2023
-
[221]
Yuansen Zhang, Xiao Wang, Zhiheng Xi, Han Xia, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024. RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions. arXiv preprint arXiv:2402.16431 (2024)
2024 arXiv
-
[222]
Yabin Zhang, Wenjie Zhu, Hui Tang, Zhiyuan Ma, Kaiyang Zhou, and Lei Zhang. 2024. Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 28718–28728. Proc. ...
2024
-
[223]
Abdelrahman Zayed, Gonçalo Mordido, Samira Shabanian, Ioana Baldini, and Sarath Chandar. 2024. Fairness-aware structured pruning in transformers. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 22484–22492
2024
-
[224]
Chongzhi Zhang, Aishan Liu, Xianglong Liu, Yitao Xu, Hang Yu, Yuqing Ma, and Tianlin Li. 2020. Interpreting and improving adversarial robustness of deep neural networks with neuron sensitivity. IEEE Transactions on Image Processing 30 (2020), 1291–1304
2020
-
[225]
Shuai Zhao, Xiaohan Wang, Linchao Zhu, and Yi Yang. 2024. Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id= kIP0duasBb
2024
-
[226]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)
2023 arXiv
-
[227]
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai man Cheung, and Min Lin. 2023. On Evaluating Adversarial Robustness of Large Vision-Language Models. InThirty-seventh Conference on Neural Information Processing Systems. https://openreview. net/forum?id=xbbknN9QFs
2023
-
[228]
Feng Zhao, Wan Xianlin, Cheng Yan, and Chu Kiong Loo. 2024. Correcting Language Model Bias for Text Classification in True Zero-Shot Learning. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING...
2024
-
[229]
Jiaxu Zhao, Meng Fang, Zijing Shi, Yitong Li, Ling Chen, and Mykola Pechenizkiy. 2023. CHBias: Bias Evaluation and Mitigation of Chinese Conversational Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...
2023
-
[230]
Rui Zheng, Wei Shen, Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Haoran Huang, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024. Improving Generalization of Alignment with Human Preferences through Group Invariant Learning. In The Twelfth International Conf...
2024
-
[231]
Li Zhong and Zilong Wang. 2024. Can LLM Replace Stack Overflow? A Study on Robustness and Reliability of Large Language Model Code Generation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21841–21849
2024
-
[232]
Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103 (2017)
2017 arXiv
-
[233]
Yilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi, Wenlin Zhang, Xiangru Tang, Boyu Mi, and Dragomir Radev. 2023. RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations. In Proceedings of the 61st Annual Meeting of the Association for ...
2023
-
[234]
Rui Zheng, Rong Bao, Qin Liu, Tao Gui, Qi Zhang, Xuan-Jing Huang, Rui Xie, and Wei Wu. 2022. Plugat: A plug and play module to defend against textual adversarial attack. In Proceedings of the 29th International Conference on Computational Linguistics . 2873–2882
2022
-
[235]
Xuhui Zhou, Yue Zhang, Leyang Cui, and Dandan Huang. 2020. Evaluating commonsense in pre-trained language models. InProceedings of the AAAI conference on artificial intelligence , Vol. 34. 9733–9740
2020
-
[236]
Yuhao Zhou, Wenxiang Chen, Rui Zheng, Zhiheng Xi, Tao Gui, Qi Zhang, and Xuan-Jing Huang. 2024. ORTicket: Let One Robust BERT Ticket Transfer across Different Tasks. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and ...
2024
-
[237]
Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan, Xiuhong Li, et al
-
[238]
Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, and Ting Zhong. 2023. Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...
2023
-
[239]
Wenjie Zhou, Qiang Wang, Mingzhou Xu, Ming Chen, and Xiangyu Duan. 2024. Revisiting the Self-Consistency Challenges in Multi-Choice Question Formats for Large Language Model Evaluation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lan...
2024
-
[240]
Shaolin Zhu, Menglong Cui, and Deyi Xiong. 2024. Towards robust in-context learning for machine translation with large language models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)....
2024
-
[243]
arXiv preprint arXiv:2404.14294 (2024)
A survey on efficient inference for large language models. arXiv preprint arXiv:2404.14294 (2024)
2024 arXiv
-
[244]
Zihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao, Jianan Ye, Wei Liu, Wei Wang, Xiaowei Huang, and Kaizhu Huang. 2024. Mathattack: Attacking large language models towards math solving ability. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 19750–19758
2024
-
[245]
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, et al. 2023. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528 (2023)
2023 arXiv
-
[2023]
arXiv preprint arXiv:2302.12095 (2023)
On the robustness of chatgpt: An adversarial and out-of-distribution perspective. arXiv preprint arXiv:2302.12095 (2023)
2023 arXiv
-
[2024]
Advances in Neural Information Processing Systems 36 (2024)
Efficient equivariant transfer learning from pretrained models. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[2025]
arXiv preprint arXiv:2502.13189 (2025)
MoBA: Mixture of Block Attention for Long-Context LLMs. arXiv preprint arXiv:2502.13189 (2025)
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.