REVIEW 5 major objections 6 minor 1 cited by
Curriculum Demonstration Selection for In-Context Learning
T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Difficulty-stratified demonstrations improve in-context learning across nine LLMs.
desk verdict A simple, cheap demonstration-selection idea with an overclaimed headline; the paper's own tables do not support 'consistently outperforms'. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the difficulty partition: the training set is divided into k strata by human-annotated complexity labels, and the retrieval function, either random or CLS-embedding similarity as in KATE, is constrained to return exactly one example per stratum. This forces the prompt to span the full difficulty range instead of concentrating on the most similar or most complex items. The shuffle at the end makes clear that ordering is not part of the mechanism.
What would settle it
Take a fixed training pool and a fixed model, and run CDS twice: once with true difficulty labels and once with difficulty labels randomly permuted across examples while keeping the same bucket sizes. If random-label CDS matches true-label CDS on the same test set, the curriculum signal is not doing the work. A second check is to replace human labels with model-computed difficulty, such as the model's own zero-shot accuracy on each candidate example, and require that CDS performance tracks the model-based partition rather than the human-based one.
Extended reading notes
Core claim
The central claim is that a balanced difficulty curriculum in the prompt, rather than the hardest or the most similar examples, elicits better in-context performance. Concretely, CDS partitions the training set into k difficulty groups using dataset metadata, such as the five MATH complexity levels, ARC grade levels, and LeetCode difficulty labels plus acceptance rates, then retrieves exactly one demonstration from each group and shuffles them. The paper reports that this recipe outperforms uniform random selection and KATE on all nine tested models, that similarity retrieval within levels beats random retrieval, that easy-to-hard ordering does not matter, and that the improvement grows with problem difficulty, from about 2% on easy MATH problems to 6% on hard ones. If correct, CDS is a lightweight, metadata-driven upgrade to few-shot prompting that needs no model training.
Load-bearing premise
The load-bearing premise is that human-annotated difficulty labels, such as MATH complexity levels, ARC grade levels, and LeetCode difficulty plus acceptance rates, track what each LLM actually finds hard, and that a flat one-example-per-level sample is more useful than the hardest or the most similar examples.
Editorial extensions
If this is right
- On MATH and ARC-Challenge, CDS outperforms both uniform random selection and KATE across Llama-2 7B/13B, Llama-3 8B, Mistral-7B, and Qwen-7B.
- On Mercury code generation, CDS improves the Pass metric over both baselines for CodeLlama-7B/13B, DeepSeek-Coder-6.7B, and StarCoder2-3B, and matches or slightly trails KATE on the Beyond efficiency metric.
- The performance advantage of CDS grows with problem difficulty, from roughly 2% on easy MATH problems to 6% on hard ones, and ARC-Challenge shows a similar but smaller trend.
- Presenting demonstrations easy-to-hard versus shuffled order makes no significant difference, so the benefit comes from the composition of the demonstration set, not its ordering.
- Using similarity-based retrieval inside each difficulty level is better than random retrieval within levels, so difficulty coverage and query relevance combine.
Reading between the lines
- Beyond the paper's claims: if these results hold up, any benchmark that ships with grade levels, contest tiers, or acceptance rates can be plugged into CDS without retraining, and datasets lacking such labels could be partitioned with model-based difficulty estimates, an extension the paper itself names as future work.
- A natural falsification experiment the paper does not run is assigning random difficulty labels while keeping the same bucket structure; if CDS still beats KATE, the measured gains would come from stratification itself rather than from the meaning of the labels.
- The pattern that hardest problems benefit most suggests CDS acts less like a curriculum in the training sense and more like a coverage regularizer that prevents the prompt from being dominated by easy, high-similarity examples; that hypothesis could be tested by comparing CDS to selecting k diverse examples purely by embedding distance.
- A general consequence the paper leaves implicit is that difficulty-stratified prompts could transfer to few-shot settings where the test distribution is unknown or changing, because coverage provides robustness that pure similarity cannot.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Curriculum Demonstration Selection (CDS), a method for in-context learning that partitions the training set by human-annotated difficulty levels and selects one demonstration per level, optionally using similarity-based retrieval within each level. The authors evaluate CDS against random selection and KATE on MATH, ARC-Challenge, and Mercury across five LLMs for math/commonsense and four code-specific LLMs, reporting mean and standard deviation over three seeds. The paper claims that CDS consistently outperforms both baselines and is especially effective on hard problems. However, the reported results do not support the 'consistently outperforms' claim: several CDS scores are within noise or below KATE, and no significance testing is provided. The experimental design also lacks a control that isolates the difficulty-stratification mechanism, and the reliance on unvalidated difficulty metadata raises concerns about the curriculum-based explanation.
Significance. The paper addresses a practical and timely question in in-context learning: how to select demonstrations to improve LLM performance. If CDS's effects are real, the method is attractive because it leverages metadata that is often available and adds little computational cost. The breadth of the evaluation (nine LLMs, three task types) is a strength, as are the ablations on retrieval function and demonstration ordering. However, the empirical evidence as presented is not sufficient to support the headline claim of consistent improvement; the missing control and lack of significance testing leave the magnitude and source of the reported gains unclear. The idea is still potentially valuable and the identified issues are addressable in a revision.
major comments (5)
- [4.2, Tables 1-2] The abstract and Section 4.2 state that CDS 'consistently outperforms' the baselines, but several cells in the paper's own tables contradict this. For example, Table 1 (Llama-3-8B ARC-c Avg) gives 81.0±0.15 for CDS versus 81.11±0.20 for KATE, and Table 2 (DeepSeek-Coder Beyond) gives 51.51±1.38 for CDS versus 52.24±0.11 for KATE; Mistral-7B Algebra in Table 1 is also lower for CDS (14.89±0.03 vs 14.90±0.30). Since all results are means over only three seeds and no paired significance test is reported, differences of 0.1-1.5 points with overlapping standard errors do not establish consistent improvement. Please provide a per-dataset and per-model significance analysis (e.g., paired bootstrap or permutation test) and revise the 'consistently outperforms' claim to match the evidence.
- [4.1.4, 4.2] The experimental design confounds the difficulty-stratification mechanism with the effect of simply diversifying the demonstration set. CDS differs from KATE not only by partitioning by difficulty but also by selecting exactly one demonstration per partition; a method that randomly samples one demonstration from each of k random partitions of the training set would control for the coverage/diversity effect. Without such a control, the gains attributed to the curriculum could be driven by any type of stratification. Please add a random-stratified baseline (one demonstration per partition, partitions defined by difficulty or by random grouping) and a KATE variant that selects one similar example per difficulty partition, to isolate the contribution of the difficulty-based curriculum.
- [3.2, 4.2.4] CDS relies on human-annotated difficulty metadata (MATH levels, ARC grade levels, LeetCode labels plus acceptance rates), but the paper does not validate that these metadata reflect the difficulty experienced by the target LLMs. This is particularly problematic for the claim in Section 4.2.4 and Figure 2 that CDS 'especially' improves hard problems, because the same metadata are used both to construct the curriculum and to define the difficulty bins in the evaluation. Please report per-item correlations between the metadata and model solve rates (or use a model-based difficulty measure as a robustness check) to rule out circularity.
- [3.2, 4.1.3] The relationship between the number of difficulty partitions and the number of demonstrations k is underspecified. Section 3.2 states that k is determined by the distribution of complexity scores, but Section 4.1.3 fixes k=5 for all experiments. For ARC-Challenge (grade levels 3-9, i.e., seven levels) and Mercury (Easy/Medium/Hard plus acceptance rates), it is not explained how exactly five partitions are formed. Please specify the binning procedure explicitly, including how grade levels or acceptance rates are mapped to five partitions, so that the method is reproducible.
- [2.2, 4.1.4] The related work section discusses In-Context Curriculum Learning (ICCL, [27]) and other curriculum-based demonstration selection methods, but the experiments compare CDS only against Uniform random and KATE. Since ICCL is the closest method that also applies curriculum principles to ICL, the absence of a comparison weakens the claim that CDS is a novel and effective curriculum-based selection method. Please add ICCL (or a faithful reimplementation) as a baseline, or justify why it is not applicable.
minor comments (6)
- [Appendix A, Tables 4-5] The prompt examples contain several LaTeX/formatting errors (e.g., 'äab¨', '\textbfOutput', 'lcm' without the operator formatting in Table 4, and a garbled character in Table 5). These should be fixed for readability.
- [Table 1] The caption 'Result on MATH and ARC-c dataset' does not clarify that the 'Avg' column refers to ARC-c accuracy while the first five columns are MATH topic accuracies. Please rename the columns or split the table to avoid ambiguity.
- [4.2.1] The statement that 'KATE underperforms the random baseline by approximately 9% with Qwen 7B' is inaccurate: the difference in Table 1 is 48.66-44.62 = 4.04 points, which is about 8% relative, not 9%. Please either give absolute differences or compute the relative change correctly.
- [Table 3] The E2H row reports values without standard deviations, while the Rand row includes them; please report errors for both rows or note that E2H is a single run.
- [General] The paper does not provide a link to code or data, which makes the exact partition and retrieval procedures hard to reproduce; consider releasing a public implementation.
- [Abstract/Keywords] The keywords 'ACM proceedings, LATEX, text tagging' appear to be template placeholders and should be replaced with actual keywords for this work.
Circularity Check
No significant circularity: CDS's difficulty stratification is an input recipe, and the claimed gains are empirical, not derived from the input metadata.
full rationale
The paper contains no derivation chain in which an output is equivalent to an input by construction. CDS partitions the training set using external human-annotated difficulty metadata (MATH complexity levels, ARC grade levels, LeetCode difficulty labels plus acceptance rates) and then selects one demonstration per partition. The property that the selected set covers a range of difficulty levels is true by definition of the algorithm, but that is the method's construction, not a circular explanation of downstream performance. No parameter is fitted to the test set or to the evaluation metric; k=5 is fixed and retrieval uses a fixed pretrained encoder. The predicted quantities (MATH/ARC-c accuracy, Mercury Pass/Beyond) are not equal to the difficulty metadata used as input. The Mercury benchmark is co-authored by overlapping authors, but it is an external evaluation set rather than a load-bearing proof, so this is at most an independence concern, not circularity. Section 6 candidly acknowledges that the difficulty-metadata assumption may not hold in other domains, which is a validity limitation rather than a circular step. The main legitimate criticism is empirical: the abstract's 'consistently outperforms' wording is not supported by several table entries (e.g., CodeLlama-7b Pass 36.98±1.93 vs KATE 36.85±0.90; DeepSeek-Coder Beyond 51.51±1.38 vs KATE 52.24±0.11; Llama-3-8B ARC-c 81.0±0.15 vs KATE 81.11±0.20), and no paired significance testing is reported. That is a correctness/statistical issue, not circularity.
Assumptions & free parameters
free parameters (2)
- number of demonstrations (k) =
5
- number of difficulty partitions =
set equal to k (5)
assumptions (3)
- domain assumption Human-annotated difficulty metadata (grade levels, acceptance rates) is a valid proxy for LLM task difficulty.
- ad hoc to paper A balanced set with one demonstration per difficulty level improves in-context learning.
- domain assumption CoT prompting with greedy decoding and exact-match extraction is a valid evaluation setup for the tested models.
Cite this review
Pith. "Pith review of Curriculum Demonstration Selection for In-Context Learning." pith.science (2026). https://pith.science/paper/AYE25WW5
@misc{pith2026241118126,
author = {Pith},
title = {Pith review of: Curriculum Demonstration Selection for In-Context Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/AYE25WW5}},
note = {Machine review of arXiv:2411.18126}
}
read the original abstract
Large Language Models (LLMs) have shown strong in-context learning (ICL) abilities with a few demonstrations. However, one critical challenge is how to select demonstrations to elicit the full potential of LLMs. In this paper, we propose Curriculum Demonstration Selection (CDS), a novel demonstration selection method for ICL. Instead of merely using similarity, CDS additionally partitions samples by their complexity measurements. Following curriculum learning, CDS then selects demonstrations from easy to difficult. Thus the selected demonstrations cover a wide range of difficulty levels, enabling LLMs to learn from varied complexities within the training set. Experiments demonstrate that our CDS consistently outperforms baseline methods, achieving notable improvements across nine LLMs on three benchmarks. Moreover, CDS proves especially effective in enhancing LLM performance in solving challenging problems.
Figures
Forward citations
Cited by 1 Pith paper
-
DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
DICE dynamically retrieves the most relevant in-context demonstrations at each agent step, and in this preprint it raises exact-match and success-rate scores on HotpotQA, ALFWorld, and Webshop across ReAct, Reflexion,...
Reference graph
Works this paper leans on
-
[27]
Yinpeng Liu, Jiawei Liu, Xiang Shi, Qikai Cheng, and Wei Lu. 2024. Let’s Learn Step by Step: Enhancing In-Context Learning Ability with Curriculum Learning. arXiv preprint arXiv:2402.10738 (2024)
arXiv 2024
-
[1]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609 (2023)
arXiv 2023
-
[2]
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. In International Conference on Machine Learning . https: //api.semanticscholar.org/CorpusID:873046
work page 2009
-
[3]
Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)
arXiv 2020
-
[4]
Zui Chen, Yezeng Chen, Jiaqi Han, Zhijie Huang, Ji Qi, and Yi Zhou. 2024. An Empirical Study of Data Ability Boundary in LLMs’ Math Reasoning. arXiv preprint arXiv:2403.00799 (2024)
arXiv 2024
-
[5]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457 (2018)
arXiv 2018
-
[6]
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022. A survey on in-context learning. arXiv preprint arXiv:2301.00234 (2022)
arXiv 2022
-
[7]
Andrew Drozdov, Honglei Zhuang, Zhuyun Dai, Zhen Qin, Razieh Rahimi, Xu- anhui Wang, Dana Alon, Mohit Iyyer, Andrew McCallum, Donald Metzler, and SAC’25, March 31 –April 4, 2025, Sicily, Italy Anh et al. Kai Hui. 2023. PaRaDe: Passage Ranking using Demonstrations with LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda B...
doi:10.18653/v1/202 2025
Show all 54 references
-
[8]
Mingzhe Du, Anh Tuan Luu, Bin Ji, and See-Kiong Ng. 2024. Mercury: An efficiency benchmark for llm code synthesis. arXiv preprint arXiv:2402.07844 (2024)
2024 arXiv
-
[9]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
2024 arXiv
-
[10]
Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. In The Eleventh Inter- national Conference on Learning Representations
2022
-
[11]
Hila Gonen, Srini Iyer, Terra Blevins, Noah Smith, and Luke Zettlemoyer. 2023. Demystifying Prompts in Language Models via Perplexity Estimation. InFindings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Associa...
2023 doi
-
[12]
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Tora: A tool-integrated reasoning agent for mathematical problem solving. arXiv preprint arXiv:2309.17452 (2023)
2023 arXiv
-
[13]
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence.arXiv preprint arXiv:2401.14196 (2024)
2024 arXiv
-
[14]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language under- standing. arXiv preprint arXiv:2009.03300 (2020)
2020 arXiv
-
[15]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874 (2021)
2021 arXiv
-
[16]
Nhat Hoang, Xuan Long Do, Duc Anh Do, Duc Anh Vu, and Anh Tuan Luu. 2024. ToXCL: A Unified Framework for Toxic Speech Detection and Explanation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language...
2024
-
[17]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al . 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[18]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei. 2020. Scaling Laws for Neural Language Models. ArXiv abs/2001.08361 (2020). https://api. semanticscholar.org/CorpusID:210861095
2020 arXiv
-
[19]
Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, and Deepak Ramachan- dran. 2023. LAMBADA: Backward Chaining for Automated Reasoning in Natural Language. In Proceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers) , An...
2023 doi
-
[20]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems 35 (2022), 22199–22213
2022
-
[21]
Ghader Kurdi, Jared Leo, Bijan Parsia, Uli Sattler, and Salam Al-Emari. 2019. A Systematic Review of Automatic Question Generation for Educational Purposes. International Journal of Artificial Intelligence in Education 30 (2019), 121–204. https://api.semanticscholar.org/Corpus...
2019
-
[22]
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. 2022. Solving quantitative reasoning problems with language models. Advances in Neural Information Processing Syste...
2022
-
[23]
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. 2023. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161 (2023)
2023 arXiv
-
[24]
Xiaonan Li and Xipeng Qiu. 2023. Finding support examples for in-context learning. arXiv preprint arXiv:2302.13539 (2023)
2023 arXiv
-
[25]
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. What Makes Good In-Context Examples for GPT-3?. In Pro- ceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowl- edge Extraction and Integration for Deep Learning ...
2022 doi
-
[26]
Yinhan Liu. 2019. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692 (2019)
2019 arXiv
-
[28]
Thong Nguyen, Yi Bin, Xiaobao Wu, Xinshuai Dong, Zhiyuan Hu, Khoi Le, Cong- Duy Nguyen, See-Kiong Ng, and Luu Anh Tuan. 2025. Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning. In European Conference on Computer Vision . Springer, 77–98
2025
-
[29]
Thong Nguyen and Anh Tuan Luu. 2021. Contrastive learning for neural topic model. Advances in neural information processing systems 34 (2021), 11974–11986
2021
-
[30]
Thong Nguyen, Anh Tuan Luu, Truc Lu, and Tho Quan. 2021. Enriching and con- trolling global semantics for text summarization. arXiv preprint arXiv:2109.10616 (2021)
2021 arXiv
-
[31]
Thong Thanh Nguyen and Anh Tuan Luu. 2022. Improving neural cross-lingual abstractive summarization via employing optimal transport distance for knowl- edge distillation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 11103–11111
2022
-
[32]
Fengjun Pan, Xiaobao Wu, Zongrui Li, and Luu Anh Tuan. 2024. Are LLMs Good Zero-Shot Fallacy Classifiers?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 14338–14364
2024
-
[33]
Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, and Preslav Nakov. 2023. Fact-Checking Complex Claims with Program-Guided Reasoning. In Annual Meeting of the Association for Compu- tational Linguistics (ACL). Association for Computational L...
2023
-
[34]
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only.arX...
2023 arXiv
-
[35]
Phuoc Van Long Pham, Anh Vu Duc, Nhat Minh Hoang, Xuan Long Do, and Anh Tuan Luu. 2024. ChatGPT as a Math Questioner? Evaluating Chat- GPT on Generating Pre-university Math Questions. In Proceedings of the 39th ACM/SIGAPP Symposium on Applied Computing (Avila, Spain) (SAC ’24)...
2024
-
[36]
Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2022. Reasoning with language model prompting: A survey. arXiv preprint arXiv:2212.09597 (2022)
2022 arXiv
-
[37]
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiao- qing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)
2023 arXiv
-
[38]
Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning To Retrieve Prompts for In-Context Learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Hu- man Language Technologies, Marine Carpuat, Ma...
2022 doi
-
[39]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Commun. ACM 64, 9 (2021), 99–106
2021
-
[40]
Dominik Sobania, Martin Briesch, Carol Hanna, and Justyna Petke. 2023. An analysis of the automatic bug fixing performance of chatgpt. In 2023 IEEE/ACM International Workshop on Automated Program Repair (APR) . IEEE, 23–30
2023
-
[41]
Taylor Sorensen, Joshua Robinson, Christopher Rytting, Alexander Shaw, Kyle Rogers, Alexia Delorey, Mahmoud Khalil, Nancy Fulda, and David Wingate. 2022. An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels. In Proceedings of the 60th Annual Meet...
2022
-
[42]
Weisong Sun, Chunrong Fang, Yudu You, Yun Miao, Yi Liu, Yuekang Li, Gelei Deng, Shenghan Huang, Yuchen Chen, Quanjun Zhang, et al. 2023. Automatic code summarization via chatgpt: How far are we?arXiv preprint arXiv:2305.12865 (2023)
2023 arXiv
-
[43]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Com- monsenseQA: A Question Answering Challenge Targeting Commonsense Knowl- edge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Hu...
2019 doi
-
[44]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)
2023 arXiv
-
[45]
Xinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers, and William Yang Wang. 2024. Large language models are latent variable models: Explaining SIG Proceedings Paper in LaTeX Format SAC’25, March 31 –April 4, 2025, Sicily, Italy and finding good demonstrations for in-context ...
2024
-
[46]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837
2022
-
[47]
Xiaobao Wu, Liangming Pan, William Yang Wang, and Luu Anh Tuan. 2024. AKEW: Assessing Knowledge Editing in the Wild. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 15118–15133
2024
-
[48]
Zhiyong Wu, Yaoxiang Wang, Jiacheng Ye, and Lingpeng Kong. 2023. Self- Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics...
2023
-
[49]
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119) , Hal D...
2020
-
[50]
Quanjun Zhang, Tongke Zhang, Juan Zhai, Chunrong Fang, Bowen Yu, Weisong Sun, and Zhenyu Chen. 2023. A critical review of large language model on software engineering: An example from chatgpt and automated program repair. arXiv preprint arXiv:2310.08879 (2023)
2023 arXiv
-
[51]
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B Hashimoto. 2024. Benchmarking large language models for news summarization. Transactions of the Association for Computational Linguistics 12 (2024), 39–57
2024
-
[52]
Yiming Zhang, Shi Feng, and Chenhao Tan. 2022. Active Example Selection for In-Context Learning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Li...
2022 doi
-
[53]
Wenting Zhao, Ye Liu, Yao Wan, Yibo Wang, Qingyang Wu, Zhongfen Deng, Jiangshu Du, Shuaiqi Liu, Yunlong Xu, and Philip Yu. 2024. 𝑘NN-ICL: Composi- tional Task-Oriented Parsing Generalization with Nearest Neighbor In-Context Learning. In Proceedings of the 2024 Conference of th...
2024
-
[54]
aaabcc". Another way is to delete one ’b’ and one ’c’, resulting in the good string
Aojun Zhou, Ke Wang, Zimu Lu, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu, Anya Jia, Linqi Song, Mingjie Zhan, et al. 2023. Solving challenging math word problems using gpt-4 code interpreter with code-based self-verification. arXiv preprint arXiv:2308.07921 (2023). A Pro...
2023 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.