REVIEW 3 major objections 6 minor 119 references
ArgInstruct: Specialized Instruction Fine-Tuning for Computational Argumentation
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A domain-specific instruction fine-tuning recipe for computational argumentation makes an LLM stronger on unseen CA tasks without sacrificing its general instruction-following ability.
desk verdict Useful CA benchmark and dataset, but the 'unseen task' claim is inflated by leaking held-out instructions into the generation seed pool. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the adaptation of the self-instruct process to a single knowledge domain: an LLM generates new CA instructions using few-shot examples drawn from the seed task pool, filters them for CA relevance and novelty (ROUGE-L F1 threshold 0.7), then creates input-output instances per task type (classification, regression, generation). Mixing the resulting 52k CA tasks with general instruction tasks at fine-tuning time is what lets the model keep its generalization abilities while gaining CA specialization.
What would settle it
Retrain the pipeline while sampling few-shot examples for instruction generation only from the 84 training seed tasks (excluding the 21 held-out tasks), then evaluate zero-shot on the held-out tasks; if the gains over the base model shrink or vanish, the reported unseen-task improvement is substantially caused by leakage rather than transferable specialization.
Extended reading notes
Core claim
Specialized instruction fine-tuning for computational argumentation significantly improves zero-shot generalization to unseen CA tasks while preserving general NLP performance. The authors construct a seed set of 105 CA tasks from 30 corpora, generate 52,445 additional CA tasks with a self-instruct process adapted to the domain, and fine-tune Gemma-2-9B on combinations of seed CA, generated CA, and general instruction data. The combination of all three sources ranks best on unseen CA tasks (mean rank 2.0), and the final ArgInstruct model outperforms comparable instruction-following models on unseen CA tasks in zero-shot evaluation (F1 .65, mean rank 2.33), while retaining its performance on the SuperNI benchmark.
Load-bearing premise
The 21 tasks used to measure 'unseen' generalization are assumed to be entirely new to the model, but the instruction-generation step samples few-shot examples from the full seed task pool including those same 21 tasks, so the generated training data may encode the structure of the test tasks.
Editorial extensions
If this is right
- CA-specialized instruction tuning yields a single model that handles argument mining, assessment, and generation tasks in zero-shot mode, without task-specific fine-tuning.
- Mixing general instruction data with domain data prevents catastrophic forgetting of general NLP abilities, as measured on SuperNI.
- The method is a general recipe: any domain with a collection of seed tasks and datasets could be turned into a specialized instruction-following model.
- For tasks with large amounts of task-specific training data, dedicated fine-tuned models still outperform the specialist, so the approach targets broad-coverage settings rather than per-task peak performance.
Reading between the lines
- If the leakage from sampling few-shot examples over the full seed pool including held-out tasks is real, the reported gains on 'unseen' tasks are an upper bound; retraining with a clean split would likely shrink the margin, though relative ordering of data mixtures may persist.
- Regression-type CA tasks remain unsolved by all tested models (all MASE scores above the mean baseline), suggesting that domain-specialized instruction tuning may need task-type-specific output decoding or training objectives for such tasks.
- A direct transfer test to another knowledge-intensive domain, such as education, would clarify whether the mechanism is domain-specialized instruction diversity or something specific to argumentation data; the paper leaves this to future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ArgInstruct, a method for specialized instruction fine-tuning in computational argumentation (CA). The authors manually craft 105 seed tasks from 30 argumentation corpora, use a self-instruct-style pipeline with Meta-Llama-3-70B to generate about 52k additional CA tasks, and fine-tune Gemma-2-9B on combinations of seed, generated, and general instruction data. They evaluate on held-out instances of training tasks, on 21 tasks reserved as 'unseen', on SuperNI, and against task-specific SOTA and several instruction-following LLMs. The central claim is that combining seedCA, genCA, and general data yields the best mean rank on unseen CA tasks (rank 2.0) while preserving general instruction-following performance.
Significance. If the unseen-task results are valid, the paper makes a useful contribution: a large CA instruction dataset, a 105-task benchmark, and evidence that domain-specialized instruction tuning can outperform general instruction-following models of comparable size. The release of code and data, the per-task results in the appendices, the manual quality evaluation of generated data, and the use of significance tests are clear strengths. However, the central generalization claim is conditional on the test tasks being excluded from the data-generation pool, and the current setup does not satisfy that condition.
major comments (3)
- [Section 4.1 and Section 5.1] The 'unseen' test tasks are not truly unseen during data generation. T0 contains all 105 seed tasks, including the 21 tasks from the nine datasets marked with '*' in Table 1 and reserved for testing in Section 5.1. Section 4.1 states that 'Based on the 105 seed tasks, we generated 52,445 additional CA tasks,' and Section 3.1 describes a self-instruct loop where each generation step samples few-shot instructions from I0 ∪ I<i, with I0 being all seed instructions. The diversity filter (ROUGE-L < 0.7) additionally keeps generated instructions within a similarity band of every existing seed instruction, including the held-out ones. Moreover, instance generation uses 'tasks from T0 representative of each task type,' so held-out tasks also serve as templates for the generated input-output instances. As a result, the genCA training data likely contains instructions and instance patterns that are paraphrases or near-variants of the held-out test tasks. The key result in Table 3(b) — that the full combination (seedCA+genCA+general) outperforms seedCA alone on unseen CA tasks (rank 2.0 vs. 5.7) — is therefore confounded: the improvement may reflect training on task-structure information derived from the test tasks rather than genuine zero-shot generalization. This issue is load-bearing for the abstract and Section 5.2. Please regenerate genCA using only the 84 training seed tasks (or otherwise exclude the held-out instructions from the generation, filtering, and instance-generation pools) and re-run Table 3 and all subsequent comparisons that depend on it.
- [Section 5.3 and Table 4] The SuperNI evaluation is under-specified, which weakens the second half of the paper's central claim that general instruction-following ability 'remains stable.' The manuscript does not state which subset of the 1600+ SuperNI tasks was used, how the aggregate F1 and ROUGE-L scores were computed, or how the mean rank in Table 4 was obtained. With only four rows and no per-task results, the reader cannot independently verify that specialization does not harm general performance. Please provide the SuperNI evaluation protocol, the list of tasks, the aggregation method, and per-task results, ideally following the standard SuperNI evaluation setup.
- [Table 3] The Wilcoxon signed-rank tests are reported only as footnote markers ('†' and '‡'), without the number of paired observations, the test direction, or effect sizes. Since the unseen-task evaluation has only 21 tasks, these details matter for interpreting the claim of 'significant improvement.' More importantly, given the contamination described in the first major comment, even a statistically significant difference between ArgInstruct and the seedCA-only baseline would not currently establish zero-shot generalization; it would only show that training on generated tasks derived from the held-out instructions improves scores on those same held-out tasks. Please report the test details and re-evaluate after fixing the leakage.
minor comments (6)
- [Appendix E, Table 9] The task name 'Same Aebate Argument' should be corrected to 'Same Debate Argument.'
- [References] The reference for Wang et al. (2023) contains a corrupted author string, 'Swaroop and/ Liu'; the correct authorship should be 'Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi.'
- [Section 7 (Limitations)] The Limitations section contains the sentence 'we hope that our analysis in Section 4, which uses the entire test sets, gives readers an idea about the transferability,' but Section 4 contains dataset statistics rather than analyses using the entire test sets; the cross-reference should point to the relevant evaluation section or be reworded.
- [Figure 2] Figure 2 is very dense and difficult to interpret because of the repeated arrows and the many small labels; a simplified numbered diagram of the pipeline would improve readability.
- [Section 5.1] The description of the 100-instance sampling says 'covering the full range for regression,' but it is not clear whether all 100 instances are used for every regression task or how tasks with fewer than 100 unique values are handled; please clarify the sampling procedure.
- [Section 5.4 and Table 5] Table 5 reports MAE, while Section 5.1 defines MASE as the regression metric; please align the metric names and definitions between the main evaluation and the SOTA comparison.
Circularity Check
Unseen-task evaluation is contaminated: the held-out test instructions are part of the seed pool used to generate the genCA training data.
-
self definitional
[Section 4.1 (CA Task Generation); Section 5.1 (Experimental Setup); Section 3.1 (Instruction Generation)]
"Based on the 105 seed tasks, we generated 52,445 additional CA tasks. ... We reserved 21 seed tasks (20% of 105) from nine CA datasets as unseen test tasks (Table 1), balancing across argument mining, assessment, and generation. ... In each generation step i, we randomly sample a subset ˜I of size l ≥ 1 from the instructions I0 ∪ I<i in T0."
The 21 'unseen' test tasks are a subset of the 105 seed tasks T0, and T0's instructions I0 form the few-shot pool for the self-instruct generator that produced the genCA training set. Therefore the held-out test instructions are inputs to the construction of the training data, so the column labeled 'Unseen CA Tasks' in Table 3 does not measure genuinely unseen tasks. The claimed zero-shot generalization to those tasks is partially self-referential: the model was trained on instructions generated from the very task instructions it is then evaluated on. The ROUGE-L diversity filter (threshold 0.7 against any existing seed instruction) further keeps generated training instructions close to the held-out seed instructions.
full rationale
The paper's main non-circular content is substantial: the CA seed benchmark is assembled from external corpora, the general-task evaluation on SuperNI is an external benchmark, and the comparison against other instruction-following LLMs is an empirical benchmark with no fitted parameter being renamed as a prediction. Self-citations to the authors' prior datasets and papers are normal related-work citations and are not load-bearing for the method's derivation. The only significant circularity issue is the construction of the 'unseen' evaluation: since the 21 held-out test tasks are part of the 105 seed tasks, and since the self-instruct generation loop samples few-shot instructions from the full seed pool I0, the genCA training data is derived from the same instructions used as the unseen test set. This makes the 'unseen CA task' generalization claim partially circular by the paper's own definitions, though the seen-task results, SuperNI stability results, and comparisons to external models remain independent evidence.
Assumptions & free parameters
assumptions (4)
- domain assumption Supervised instruction fine-tuning on a diverse set of tasks transfers to unseen tasks of the same domain.
- domain assumption The 105 seed tasks, derived from 30 datasets, are representative of the space of computational argumentation tasks.
- domain assumption Automatic metrics (micro-F1, MASE, ROUGE-L) adequately measure performance on classification, regression, and generation CA tasks.
- domain assumption The self-instruct generation process produces valid CA tasks with correct instances.
Cite this review
Pith. "Pith review of ArgInstruct: Specialized Instruction Fine-Tuning for Computational Argumentation." pith.science (2026). https://pith.science/paper/7CXKZTJC
@misc{pith2026250522076,
author = {Pith},
title = {Pith review of: ArgInstruct: Specialized Instruction Fine-Tuning for Computational Argumentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/7CXKZTJC}},
note = {Machine review of arXiv:2505.22076}
}
read the original abstract
Training large language models (LLMs) to follow instructions has significantly enhanced their ability to tackle unseen tasks. However, despite their strong generalization capabilities, instruction-following LLMs encounter difficulties when dealing with tasks that require domain knowledge. This work introduces a specialized instruction fine-tuning for the domain of computational argumentation (CA). The goal is to enable an LLM to effectively tackle any unseen CA tasks while preserving its generalization capabilities. Reviewing existing CA research, we crafted natural language instructions for 105 CA tasks to this end. On this basis, we developed a CA-specific benchmark for LLMs that allows for a comprehensive evaluation of LLMs' capabilities in solving various CA tasks. We synthesized 52k CA-related instructions, adapting the self-instruct process to train a CA-specialized instruction-following LLM. Our experiments suggest that CA-specialized instruction fine-tuning significantly enhances the LLM on both seen and unseen CA tasks. At the same time, performance on the general NLP tasks of the SuperNI benchmark remains stable.
Figures
Reference graph
Works this paper leans on
-
[1]
Rob Abbott, Brian Ecker, Pranav Anand, and Marilyn Walker. 2016. https://aclanthology.org/L16-1704 I nternet argument corpus 2.0: An SQL schema for dialogic social media and the corpora to go with it . In Proceedings of the Tenth International Conference on Language Resources and Evaluation ( LREC '16) , pages 4445--4452, Portoro z , Slovenia. European La...
2016
-
[2]
Yamen Ajjour, Milad Alshomary, Henning Wachsmuth, and Benno Stein. 2019. https://doi.org/10.18653/v1/D19-1290 Modeling frames in argumentation . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2922--2932, Hong Kong, Chi...
-
[3]
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. https://doi.org/10.1145/3292500.3330701 Optuna: A next-generation hyperparameter optimization framework . In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD '19, page 2623–2631, New York, NY, USA. Association for Comp...
arXiv 2019
-
[4]
Khalid Al-Khatib, Henning Wachsmuth, Matthias Hagen, Jonas K \"o hler, and Benno Stein. 2016 a . https://doi.org/10.18653/v1/N16-1165 Cross-domain mining of argumentative text through distant supervision . In Proceedings of the 2016 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies , pa...
-
[5]
Khalid Al-Khatib, Henning Wachsmuth, Johannes Kiesel, Matthias Hagen, and Benno Stein. 2016 b . https://aclanthology.org/C16-1324/ A news editorial corpus for mining argumentation strategies . In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers , pages 3433--3443, Osaka, Japan. The COLING 2016 Or...
2016
-
[6]
Tariq Alhindi and Debanjan Ghosh. 2021. https://aclanthology.org/2021.bea-1.22/ Sharks are not the threat humans are : Argument component segmentation in school student essays . In Proceedings of the 16th Workshop on Innovative Use of NLP for Building Educational Applications, pages 210--222, Online. Association for Computational Linguistics
2021
-
[7]
Milad Alshomary, Wei-Fan Chen, Timon Gurcke, and Henning Wachsmuth. 2021. https://doi.org/10.18653/v1/2021.eacl-main.17 Belief-based generation of argumentative claims . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 224--233, Online. Association for Computational Linguistics
-
[8]
ca.\ 350 B.C.E
Aristotle. ca.\ 350 B.C.E. / translated 2007. On Rhetoric: A Theory of Civic Discourse . Oxford University Press, Oxford, UK. Translated by George A. Kennedy
2007
Show all 119 references
-
[9]
Jianzhu Bao, Bojun Jin, Yang Sun, Yice Zhang, Yuhang He, and Ruifeng Xu. 2024. https://doi.org/10.3390/electronics13204088 A comparison-based framework for argument quality assessment . Electronics, 13(20)
2024 doi
-
[10]
Roy Bar-Haim, Indrajit Bhattacharya, Francesco Dinuzzo, Amrita Saha, and Noam Slonim. 2017. https://aclanthology.org/E17-1024/ Stance classification of context-dependent claims . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computationa...
2017
-
[11]
Tilman Beck, Ji-Ung Lee, Christina Viehmann, Marcus Maurer, Oliver Quiring, and Iryna Gurevych. 2021. https://doi.org/10.18653/v1/2021.acl-long.1 Investigating label suggestions for opinion mining in G erman covid-19 social media . In Proceedings of the 59th Annual Meeting of ...
2021 doi
-
[12]
Filip Boltu z i \'c and Jan S najder. 2014. https://doi.org/10.3115/v1/W14-2107 Back up your stance: Recognizing arguments in online discussions . In Proceedings of the First Workshop on Argumentation Mining, pages 49--58, Baltimore, Maryland. Association for Computational Linguistics
2014 doi
- [13]
-
[14]
J \'e r \'e mie Cabessa, Hugo Hernault, and Umer Mushtaq. 2025. https://aclanthology.org/2025.coling-main.442/ Argument mining with fine-tuned large language models . In Proceedings of the 31st International Conference on Computational Linguistics, pages 6624--6635, Abu Dhabi,...
2025
-
[15]
Cayque Monteiro Castro Nascimento and André Silva Pimentel. 2023. https://doi.org/10.1021/acs.jcim.3c00285 Do large language models understand chemistry? a conversation with chatgpt . Journal of Chemical Information and Modeling, 63(6):1649--1655
2023 doi
-
[16]
Guizhen Chen, Liying Cheng, Anh Tuan Luu, and Lidong Bing. 2024 a . https://aclanthology.org/2024.acl-long.126 Exploring the potential of large language models in computational argumentation . In Proceedings of the 62nd Annual Meeting of the Association for Computational Lingu...
2024
-
[17]
Zaiqian Chen, Daniel Verdi do Amarante, Jenna Donaldson, Yohan Jo, and Joonsuk Park. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.609 Argument mining for review helpfulness prediction . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Process...
2022 doi
-
[18]
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. 2024 b . https://dl.acm.org/doi/10.5555/3692070.3692326 Self-play fine-tuning converts weak language models to strong language models . In Proceedings of the 41st International Conference on Machine Learning, ...
2024
-
[19]
Yew Ken Chia, Pengfei Hong, Lidong Bing, and Soujanya Poria. 2024. https://aclanthology.org/2024.scalellm-1.4 I nstruct E val: Towards holistic evaluation of instruction-tuned large language models . In Proceedings of the First edition of the Workshop on the Scaling Behavior o...
2024
-
[20]
Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios N. Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I. Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. https://dl.acm.org/doi/abs/10.5555/3692070.3692401 Chatbot arena: an open platform for evaluating ...
2024
-
[21]
Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tai, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinso...
2024
-
[22]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[23]
Yann Dubois, Percy Liang, and Tatsunori Hashimoto. 2024. https://openreview.net/forum?id=CybBmzWBX0 Length-controlled alpacaeval: A simple debiasing of automatic evaluators . In First Conference on Language Modeling
2024
-
[24]
Judith Eckle-Kohler, Roland Kluge, and Iryna Gurevych. 2015. https://doi.org/10.18653/v1/D15-1267 On the role of discourse markers for discriminating claims and premises in argumentative discourse . In Proceedings of the 2015 Conference on Empirical Methods in Natural Language...
2015 doi
-
[25]
Lilach Eden, Yoav Kantor, Matan Orbach, Yoav Katz, Noam Slonim, and Roy Bar-Haim. 2023. https://doi.org/10.18653/v1/2023.emnlp-industry.46 Welcome to the real world: Efficient, incremental and scalable key point analysis . In Proceedings of the 2023 Conference on Empirical Met...
2023 doi
-
[26]
Liat Ein - Dor, Eyal Shnarch, Lena Dankin, Alon Halfon, Benjamin Sznajder, Ariel Gera, Carlos Alzate, Martin Gleize, Leshem Choshen, Yufang Hou, Yonatan Bilu, Ranit Aharonov, and Noam Slonim. 2020. https://doi.org/10.1609/AAAI.V34I05.6270 Corpus wide argument mining - A workin...
2020 doi
-
[27]
Mohamed Elaraby, Diane Litman, Xiang Lorraine Li, and Ahmed Magooda. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.836 Persuasiveness of generated free-text rationales in subjective decisions: A case study on pairwise argument ranking . In Findings of the Association f...
2024 doi
-
[28]
Marc Feger and Stefan Dietze. 2024. https://aclanthology.org/2024.lrec-main.1349/ TACO -- T witter arguments from CO nversations . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), page...
2024
-
[29]
Roni Friedman, Lena Dankin, Yufang Hou, Ranit Aharonov, Yoav Katz, and Noam Slonim. 2021. https://doi.org/10.18653/v1/2021.argmining-1.16 Overview of the 2021 key point analysis shared task . In Proceedings of the 8th Workshop on Argument Mining, pages 154--164, Punta Cana, Do...
2021 doi
-
[30]
Gemma Team , Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le La...
2024 arXiv
-
[31]
Martin Gleize, Eyal Shnarch, Leshem Choshen, Lena Dankin, Guy Moshkowich, Ranit Aharonov, and Noam Slonim. 2019. https://doi.org/10.18653/v1/P19-1093 Are you convinced? choosing the more convincing evidence with a S iamese network . In Proceedings of the 57th Annual Meeting of...
2019 doi
-
[32]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[33]
Shai Gretz, Roni Friedman, Edo Cohen-Karlik, Assaf Toledo, Dan Lahav, Ranit Aharonov, and Noam Slonim. 2020. A large-scale dataset for argument quality ranking: Construction and analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7805--7813
2020
-
[34]
Giulia Grundler, Piera Santin, Andrea Galassi, Federico Galli, Francesco Godano, Francesca Lagioia, Elena Palmieri, Federico Ruggeri, Giovanni Sartor, and Paolo Torroni. 2022. https://aclanthology.org/2022.argmining-1.14/ Detecting arguments in CJEU decisions on fiscal state a...
2022
-
[35]
Ivan Habernal, Judith Eckle-Kohler, and Iryna Gurevych. 2014. http://tubiblio.ulb.tu-darmstadt.de/104647/ Argumentation mining on the web from information seeking perspective . In Proceedings of the Workshop on Frontiers and Connections between Argumentation Theory and Natural...
2014
-
[36]
Ivan Habernal and Iryna Gurevych. 2016 a . https://doi.org/10.18653/v1/D16-1129 What makes a convincing argument? empirical analysis and detecting attributes of convincingness in web argumentation . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language...
2016 doi
-
[37]
Ivan Habernal and Iryna Gurevych. 2016 b . https://doi.org/10.18653/v1/P16-1150 Which argument is more convincing? analyzing and predicting convincingness of web arguments using bidirectional LSTM . In Proceedings of the 54th Annual Meeting of the Association for Computational...
2016 doi
-
[38]
Ivan Habernal and Iryna Gurevych. 2017. https://doi.org/10.1162/COLI_a_00276 Argumentation mining in user-generated web discourse . Computational Linguistics, 43(1):125--179
2017 doi
-
[39]
Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, and Benno Stein. 2018 a . https://doi.org/10.18653/v1/N18-1175 The argument reasoning comprehension task: Identification and reconstruction of implicit warrants . In Proceedings of the 2018 Conference of the North A merican Cha...
2018 doi
-
[40]
Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, and Benno Stein. 2018 b . https://doi.org/10.18653/v1/N18-1036 Before name-calling: Dynamics and triggers of ad hominem fallacies in web argumentation . In Proceedings of the 2018 Conference of the North A merican Chapter of th...
2018 doi
-
[41]
Shohreh Haddadan, Elena Cabrio, and Serena Villata. 2019. https://doi.org/10.18653/v1/P19-1463 Yes, we can! mining arguments in 50 years of US presidential campaign debates . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4684...
2019 doi
-
[42]
Kazi Saidul Hasan and Vincent Ng. 2014. https://doi.org/10.3115/v1/D14-1083 Why are you taking this stance? identifying and classifying reasons in ideological debates . In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages ...
2014 doi
-
[43]
Annette Hautli-Janisz, Zlata Kikteva, Wassiliki Siskou, Kamila Gorska, Ray Becker, and Chris Reed. 2022. https://aclanthology.org/2022.lrec-1.352 QT 30: A corpus of argument and conflict in broadcast debate . In Proceedings of the Thirteenth Language Resources and Evaluation C...
2022
-
[44]
Philipp Heinisch, Anette Frank, Juri Opitz, Moritz Plenz, and Philipp Cimiano. 2022. https://aclanthology.org/2022.argmining-1.7 Overview of the 2022 validity and novelty prediction shared task . In Proceedings of the 9th Workshop on Argument Mining, pages 84--94, Online and i...
2022
-
[45]
Christopher Hidey, Elena Musi, Alyssa Hwang, Smaranda Muresan, and Kathy McKeown. 2017. https://doi.org/10.18653/v1/W17-5102 Analyzing the semantic types of claims and premises in an online persuasive forum . In Proceedings of the 4th Workshop on Argument Mining, pages 11--21,...
2017 doi
-
[46]
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2023. https://doi.org/10.18653/v1/2023.acl-long.806 Unnatural instructions: Tuning language models with (almost) no human labor . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics...
2023 doi
-
[47]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. https://arxiv.org/abs/2106.09685 Lora: Low-rank adaptation of large language models . Preprint, arXiv:2106.09685
2021 arXiv
-
[48]
Hyndman and George Athanasopoulos
Rob J. Hyndman and George Athanasopoulos. 2021. https://OTexts.com/fpp3 Forecasting: Principles and Practice , 3rd edition. OTexts, Melbourne, Australia
2021
-
[49]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[50]
Yohan Jo, Seojin Bang, Emaad Manzoor, Eduard Hovy, and Chris Reed. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.1 Detecting attackable sentences in arguments . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1--23, ...
2020 doi
-
[51]
Nikita Kitaev, Steven Cao, and Dan Klein. 2019. https://doi.org/10.18653/v1/P19-1340 Multilingual constituency parsing with self-attention and pre-training . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3499--3505, Florence,...
2019 doi
-
[52]
Ilia Kuznetsov, Jan Buchmann, Max Eichler, and Iryna Gurevych. 2022. Revise and resubmit: An intertextual model of text-based collaboration in peer review. Computational Linguistics, 48(4):949--986
2022
-
[53]
Anne Lauscher, Goran Glava s , and Simone Paolo Ponzetto. 2018. https://doi.org/10.18653/v1/W18-5206 An argument-annotated corpus of scientific publications . In Proceedings of the 5th Workshop on Argument Mining, pages 40--46, Brussels, Belgium. Association for Computational ...
2018 doi
-
[54]
Anne Lauscher, Henning Wachsmuth, Iryna Gurevych, and Goran Glava s . 2022. https://doi.org/10.1162/tacl_a_00525 Scientia potentia E st --- O n the role of knowledge in computational argumentation . Transactions of the Association for Computational Linguistics, 10:1392--1422
2022 doi
-
[55]
Augustin Lecler, Loïc Duron, and Philippe Soyer. 2023. https://doi.org/10.1016/j.diii.2023.02.003 Revolutionizing radiology with gpt-based models: Current applications, future possibilities and limitations of chatgpt . Diagnostic and Interventional Imaging, 104(6):269--274
2023 doi
-
[56]
Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, and Jing Xiao. 2024. https://doi.org/10.18653/v1/2024.naacl-long.421 From quantity to quality: Boosting LLM performance with self-guided data selection for instruction tuning . I...
2024 doi
-
[57]
Matthias Liebeck, Katharina Esau, and Stefan Conrad. 2016. https://doi.org/10.18653/v1/W16-2817 What to do with an airport? mining arguments in the G erman online participation project tempelhofer feld . In Proceedings of the Third Workshop on Argument Mining ( A rg M ining201...
2016 doi
-
[58]
Le, Barret Zoph, Jason Wei, and Adam Roberts
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. 2023. The flan collection: designing data and methods for effective instruction tuning. In Proceedings of the 40th International Conference ...
2023
-
[59]
Tobias Mayer, Elena Cabrio, and Serena Villata. 2020. https://hal.science/hal-02879293 Transformer-based Argument Mining for Healthcare Applications . In ECAI 2020 - 24th European Conference on Artificial Intelligence , Santiago de Compostela / Online, Spain
2020
-
[60]
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. https://doi.org/10.18653/v1/2022.acl-long.244 Cross-task generalization via natural language crowdsourcing instructions . In Proceedings of the 60th Annual Meeting of the Association for Computationa...
2022 doi
-
[61]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward ...
2023 doi
-
[62]
Nathan Ong, Diane Litman, and Alexandra Brusilovsky. 2014. https://doi.org/10.3115/v1/W14-2104 Ontology-based argument mining and automatic essay scoring . In Proceedings of the First Workshop on Argumentation Mining, pages 24--28, Baltimore, Maryland. Association for Computat...
2014 doi
-
[63]
OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, Alex Pai...
2024 arXiv
-
[64]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022
-
[65]
Joonsuk Park and Claire Cardie. 2014. http://www.aclweb.org/anthology/W/W14/W14-2105 Identifying appropriate support for propositions in online user comments . In Proceedings of the First Workshop on Argumentation Mining, pages 29--38, Baltimore, Maryland. Association for Comp...
2014
-
[66]
Joonsuk Park and Claire Cardie. 2018. https://aclanthology.org/L18-1257/ A corpus of e R ulemaking user comments for measuring evaluability of arguments . In Proceedings of the Eleventh International Conference on Language Resources and Evaluation ( LREC 2018) , Miyazaki, Japa...
2018
-
[67]
Andreas Peldszus and Manfred Stede. 2015. An annotated corpus of argumentative microtexts. In Argumentation and Reasoned Action: Proceedings of the 1st European Conference on Argumentation, Lisbon, volume 2, pages 801--815
2015
-
[68]
Isaac Persing, Alan Davis, and Vincent Ng. 2010. https://aclanthology.org/D10-1023/ Modeling organization in student essays . In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pages 229--239, Cambridge, MA. Association for Computational...
2010
-
[69]
Isaac Persing and Vincent Ng. 2013. https://aclanthology.org/P13-1026/ Modeling thesis clarity in student essays . In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 260--269, Sofia, Bulgaria. Association f...
2013
-
[70]
Isaac Persing and Vincent Ng. 2014. https://doi.org/10.3115/v1/P14-1144 Modeling prompt adherence in student essays . In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1534--1543, Baltimore, Maryland. Asso...
2014 doi
-
[71]
Isaac Persing and Vincent Ng. 2015. https://doi.org/10.3115/v1/P15-1053 Modeling argument strength in student essays . In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Proc...
2015 doi
-
[72]
Isaac Persing and Vincent Ng. 2016. https://doi.org/10.18653/v1/P16-1205 Modeling stance in student essays . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2174--2184, Berlin, Germany. Association for C...
2016 doi
-
[73]
Prakash Poudyal, Jaromir Savelka, Aagje Ieven, Marie Francine Moens, Teresa Goncalves, and Paulo Quaresma. 2020. https://aclanthology.org/2020.argmining-1.8 ECHR : Legal corpus for argument mining . In Proceedings of the 7th Workshop on Argument Mining, pages 67--75, Online. A...
2020
-
[74]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9
2019
-
[75]
Nils Reimers, Benjamin Schiller, Tilman Beck, Johannes Daxenberger, Christian Stab, and Iryna Gurevych. 2019. https://doi.org/10.18653/v1/P19-1054 Classification and clustering of arguments with contextualized word embeddings . In Proceedings of the 57th Annual Meeting of the ...
2019 doi
-
[76]
Paula Rescala, Manoel Horta Ribeiro, Tiancheng Hu, and Robert West. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.515 Can language models recognize convincing arguments? In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 8826--8837, Miami, ...
2024 doi
-
[77]
Khapra, Ehud Aharoni, and Noam Slonim
Ruty Rinott, Lena Dankin, Carlos Alzate Perez, Mitesh M. Khapra, Ehud Aharoni, and Noam Slonim. 2015. https://doi.org/10.18653/v1/D15-1050 Show me your evidence - an automatic method for context dependent evidence detection . In Proceedings of the 2015 Conference on Empirical ...
2015 doi
-
[78]
Julia Romberg. 2022. https://aclanthology.org/2022.argmining-1.11 Is your perspective also my perspective? enriching prediction with subjectivity . In Proceedings of the 9th Workshop on Argument Mining, pages 115--125, Online and in Gyeongju, Republic of Korea. International C...
2022
-
[79]
Allen Roush and Arvind Balaji. 2020. https://aclanthology.org/2020.argmining-1.1 D ebate S um: A large-scale argument mining and summarization dataset . In Proceedings of the 7th Workshop on Argument Mining, pages 1--7, Online. Association for Computational Linguistics
2020
-
[80]
Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M. Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V....
2021 arXiv
-
[81]
Nils-Jonathan Schaller, Andrea Horbach, Lars Ingver H \"o ft, Yuning Ding, Jan Luca Bahr, Jennifer Meyer, and Thorben Jansen. 2024. https://aclanthology.org/2024.lrec-main.389/ DARIUS : A comprehensive learner corpus for argument mining in G erman-language essays . In Proceedi...
2024
-
[82]
Benjamin Schiller, Johannes Daxenberger, and Iryna Gurevych. 2021. https://doi.org/10.18653/v1/2021.naacl-main.34 Aspect-controlled neural argument generation . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics...
2021 doi
-
[83]
Eyal Shnarch, Carlos Alzate, Lena Dankin, Martin Gleize, Yufang Hou, Leshem Choshen, Ranit Aharonov, and Noam Slonim. 2018. https://doi.org/10.18653/v1/P18-2095 Will it blend? blending weak and strong labeled data in a neural network for argumentation mining . In Proceedings o...
2018 doi
-
[84]
Eyal Shnarch, Leshem Choshen, Guy Moshkowich, Ranit Aharonov, and Noam Slonim. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.243 Unsupervised expressive rules provide explainability and assist human experts grasping new domains . In Findings of the Association for Comp...
2020 doi
-
[85]
Maria Skeppstedt, Andreas Peldszus, and Manfred Stede. 2018. https://doi.org/10.18653/v1/W18-5218 More or less controlled elicitation of argumentative text: Enlarging a microtext corpus via crowdsourcing . In Proceedings of the 5th Workshop on Argument Mining, pages 155--163, ...
2018 doi
-
[86]
Gabriella Skitalinskaya, Jonas Klaff, and Henning Wachsmuth. 2021. https://doi.org/10.18653/v1/2021.eacl-main.147 Learning from revisions: Quality assessment of claims in argumentation at scale . In Proceedings of the 16th Conference of the European Chapter of the Association ...
2021 doi
-
[87]
Parinaz Sobhani, Diana Inkpen, and Stan Matwin. 2015. https://doi.org/10.3115/v1/W15-0509 From argumentation mining to stance classification . In Proceedings of the 2nd Workshop on Argumentation Mining, pages 67--77, Denver, CO. Association for Computational Linguistics
2015 doi
-
[88]
Christian Stab and Iryna Gurevych. 2017 a . https://doi.org/10.1162/COLI_a_00295 Parsing argumentation structures in persuasive essays . Computational Linguistics, 43(3):619--659
2017 doi
-
[89]
Christian Stab and Iryna Gurevych. 2017 b . https://aclanthology.org/E17-1092/ Recognizing insufficiently supported arguments in argumentative essays . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational Linguistics: Volume 1, Lo...
2017
-
[90]
Christian Stab, Tristan Miller, Benjamin Schiller, Pranav Rai, and Iryna Gurevych. 2018. https://doi.org/10.18653/v1/D18-1402 Cross-topic argument mining from heterogeneous sources . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pag...
2018 doi
-
[91]
Maja Stahl, Nick D \"u sterhus, Mei-Hua Chen, and Henning Wachsmuth. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.312 Mind the gap: Automated corpus creation for enthymeme detection and reconstruction in learner arguments . In Findings of the Association for Computati...
2023 doi
-
[92]
Maja Stahl, Nadine Michel, Sebastian Kilsbach, Julian Schmidtke, Sara Rezat, and Henning Wachsmuth. 2024. https://doi.org/10.18653/v1/2024.naacl-long.145 A school student essay corpus for analyzing interactions of argumentative structure and quality . In Proceedings of the 202...
2024 doi
-
[93]
Manfred Stede and Jodi Schneider. 2018. Argumentation Mining. Number 40 in Synthesis Lectures on Human Language Technologies. Morgan & Claypool
2018
-
[94]
Benno Stein, Yamen Ajjour, Roxanne El Baff, Khalid Al-Khatib, Philipp Cimiano, and Henning Wachsmuth. 2021. Same side stance classification. In Same Side Shared Task 2019: Same Side Stance Classification Shared Task 2019, pages 1--7. CEUR Workshop Proceedings (CEUR-WS. org)
2021
-
[95]
Shahbaz Syed, Khalid Al Khatib, Milad Alshomary, Henning Wachsmuth, and Martin Potthast. 2021. https://doi.org/10.18653/v1/2021.findings-acl.306 Generating informative conclusions for argumentative texts . In Findings of the Association for Computational Linguistics: ACL-IJCNL...
2021 doi
-
[96]
Hashimoto
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca
2023
-
[97]
Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, and Noam Slonim. 2019. https://doi.org/10.18653/v1/D19-1564 Automatic argument quality assessment - new datasets and methods . In Proceedings of the 2019 Confere...
2019 doi
-
[98]
Orith Toledo-Ronen, Matan Orbach, Yonatan Bilu, Artem Spector, and Noam Slonim. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.29 Multilingual argument mining: Datasets and analysis . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 303--3...
2020 doi
-
[99]
Dietrich Trautmann. 2020. https://aclanthology.org/2020.argmining-1.5/ Aspect-based argument mining . In Proceedings of the 7th Workshop on Argument Mining, pages 41--52, Online. Association for Computational Linguistics
2020
-
[100]
Dietrich Trautmann, Johannes Daxenberger, Christian Stab, Hinrich Schütze, and Iryna Gurevych. 2020. https://doi.org/10.1609/aaai.v34i05.6438 Fine-grained argument unit recognition and classification . Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):9048--9056
2020 doi
-
[101]
Jannis Vamvas and Rico Sennrich. 2020. https://arxiv.org/abs/2003.08385 X-stance: A multilingual multi-target dataset for stance detection . CoRR, arXiv:2003.08385
2020 arXiv
-
[102]
Jacky Visser, John Lawrence, Jean Wagemans, and Chris Reed. 2019. An annotated corpus of argument schemes in us election debates. In Proceedings of the 9th Conference of the International Society for the Study of Argumentation (ISSA), 3-6 July 2018, pages 1101--1111
2019
-
[103]
Henning Wachsmuth, Gabriella Lapesa, Elena Cabrio, Anne Lauscher, Joonsuk Park, Eva Maria Vecchi, Serena Villata, and Timon Ziegenbein. 2024. https://aclanthology.org/2024.lrec-main.135 Argument quality assessment in the age of instruction-following large language models . In ...
2024
-
[104]
Henning Wachsmuth, Nona Naderi, Yufang Hou, Yonatan Bilu, Vinodkumar Prabhakaran, Tim Alberdingk Thijm, Graeme Hirst, and Benno Stein. 2017. https://aclanthology.org/E17-1017 Computational argumentation quality assessment in natural language . In Proceedings of the 15th Confer...
2017
-
[105]
Henning Wachsmuth, Manfred Stede, Roxanne El Baff, Khalid Al-Khatib, Maria Skeppstedt, and Benno Stein. 2018 a . https://aclanthology.org/C18-1318 Argumentation synthesis following rhetorical strategies . In Proceedings of the 27th International Conference on Computational Lin...
2018
-
[106]
Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018 b . https://doi.org/10.18653/v1/P18-1023 Retrieval of the best counterargument without prior topic knowledge . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pape...
2018 doi
-
[107]
Andreas Waldis, Yufang Hou, and Iryna Gurevych. 2024. https://doi.org/10.18653/v1/2024.acl-long.795 How to handle different types of out-of-distribution scenarios in computational argumentation? a comprehensive and fine-grained field study . In Proceedings of the 62nd Annual M...
2024 doi
-
[108]
Marilyn Walker, Jean Fox Tree, Pranav Anand, Rob Abbott, and Joseph King. 2012. https://aclanthology.org/L12-1643/ A corpus for research on deliberation and debate . In Proceedings of the Eighth International Conference on Language Resources and Evaluation ( LREC `12) , pages ...
2012
-
[109]
Thiemo Wambsganss, Christina Niklaus, Matthias S \"o llner, Siegfried Handschuh, and Jan Marco Leimeister. 2020. https://doi.org/10.18653/v1/2020.coling-main.74 A corpus for argumentative writing support in G erman . In Proceedings of the 28th International Conference on Compu...
2020 doi
-
[110]
Jiahao Wang, Bolin Zhang, Qianlong Du, Jiajun Zhang, and Dianhui Chu. 2024 a . https://arxiv.org/abs/2402.05123 A survey on data selection for llm instruction tuning . Preprint, arXiv:2402.05123
2024 arXiv
-
[111]
Shaokang Wang, Fuhui Sun, Xiaoyan Wang, and Li Pan. 2024 b . https://doi.org/10.1016/j.ipm.2024.103815 Incorporating target-aware knowledge into prompt-tuning for few-shot stance detection . Information Processing & Management, 61(5):103815
2024
-
[112]
Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Yizhong Wang, Yeganeh Kordi, Alisa Mishra, Swaroop and/ Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annua...
2023 doi
-
[113]
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuzni...
2022
-
[114]
Brandon T Willard and R \'e mi Louf. 2023. https://arxiv.org/abs/2307.09702 Efficient guided generation for llms
2023 arXiv
-
[115]
Fangkai Yang, Pu Zhao, Zezhong Wang, Lu Wang, Bo Qiao, Jue Zhang, Mohit Garg, Qingwei Lin, Saravan Rajmohan, and Dongmei Zhang. 2023. https://doi.org/10.18653/v1/2023.emnlp-industry.29 Empower large language model to perform better on industrial domain-specific question answer...
2023 doi
-
[116]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. In Proceedings of the 37th Internat...
2023
-
[117]
Timon Ziegenbein, Shahbaz Syed, Felix Lange, Martin Potthast, and Henning Wachsmuth. 2023. https://doi.org/10.18653/v1/2023.acl-long.238 Modeling appropriate language in argumentation . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics ...
2023 doi
-
[118]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[119]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.