REVIEW 4 major objections 6 minor 2 cited by
Large Language Models Can Self-Improve in Long-context Reasoning
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Language models can improve their own long-context reasoning by learning from the most self-consistent of their sampled answers, with no human or expert-model labels.
desk verdict A legitimate first application of MBR-style consensus self-training to long-context multi-hop QA, with consistent gains but a missing control that leaves the headline attribution under-supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the MBR score $s(y) \approx (1/N) \sum_{i=1}^{N} u(y, y_i)$, where $u(y,y')$ is the inner-product similarity between sentence embeddings of two sampled outputs, computed with jina-embeddings-v3. It turns the intuition that correct reasoning is semantically consistent into a label-free ranking of sampled trajectories; the highest-scoring output becomes the SFT target or the chosen response in ORPO, and a randomly selected low-scoring output becomes the rejected response. The other supporting piece is plan-and-solve prompting, which elicits structured step-by-step trajectories from the model before scoring.
What would settle it
Construct a long-context multi-hop question where a majority of sampled outputs contain the same wrong answer and only a minority contain the correct one; if MBR decoding selects the wrong majority output more often than the oracle and fine-tuning on those selections lowers held-out accuracy compared with random-sample fine-tuning, the consensus-as-correctness premise is falsified.
Extended reading notes
Core claim
The paper's central claim is that consensus among a model's own sampled outputs is a usable correctness signal for long-context reasoning, and that fine-tuning on the consensus-selected output converts that signal into durable capability. Concretely, SEALONG samples $N=32$ plan-and-solve reasoning trajectories per question, scores each with Minimum Bayes Risk using sentence-embedding similarity as the utility, and either supervises fine-tuning with the top-scoring output or applies ORPO with a high-scoring output preferred over a randomly chosen low-scoring one. The method raises the average SubEM of Llama-3.1-8B-Instruct from 50.8 to 55.0 across Qasper, MultiFieldQA-En, HotpotQA, MuSiQue, and 2WikiMQA, and enables Qwen-2.5-14B-Instruct to reach 54.7, exceeding the 53.1 of Qwen-2.5-32B-Instruct. It also outperforms fine-tuning on existing human- or expert-model-annotated long-context datasets at matched 2K-example budgets, indicating the self-supervision itself, not a larger or higher-quality external dataset, drives the gain.
Load-bearing premise
The method works only if the most semantically consistent output among a model's sampled answers is usually the correct one; if the model confidently repeats a wrong answer across many samples, consensus scoring will pick that wrong answer and fine-tuning will reinforce it.
Editorial extensions
If this is right
- A model trained with SEALONG improves on long-context multi-hop QA across held-out tasks it never saw during data synthesis, so the benefit is not task-specific memorization.
- SEALONG improves Qwen-2.5-7B-Instruct enough to close the gap with its 14B version, and makes 14B surpass 32B, showing self-improvement can substitute for some parameter scale.
- The gain is not an artifact of longer outputs: average output token counts stay nearly unchanged, and short-context performance on six Open LLM Leaderboard tasks is flat.
- At equal 2K-example budgets, SEALONG beats training on existing long-context datasets such as LongAlign, LongReward, and GPT-4o-MuSiQue, so the self-generated data is at least as useful as current human- or expert-annotated data.
- Performance saturates around 1K synthetic examples and 32 samples per question, suggesting the method unlocks latent capability rather than teaching a new skill that needs more data.
Reading between the lines
- If consensus-tracks-correctness holds at larger scales, self-improvement could compound: each generation of a model family could bootstrap its own long-context training data, removing the need to wait for a stronger teacher, which would change how long-context capability is scaled.
- The paper's own oracle-versus-MBR gap suggests the scoring function, not the sampling budget, is the current bottleneck; a correctness-aware utility that verifies factual claims against the context rather than only semantic overlap might close that gap.
- Because all synthetic training data comes from MuSiQue-style multi-hop questions, the method's scope is currently limited to that question type; applying the same loop to full-context reasoning, code, or math prompts would test whether the mechanism generalizes beyond multi-hop QA.
- The comparison against expert-annotated datasets is at matched 2K examples; at larger data budgets the relative advantage could shift, so the claim of superiority over human/expert data should be read as specific to this scale and setup.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SEALONG, a self-improvement method for long-context reasoning. For each training question, the method samples N outputs from the base model, scores them with Minimum Bayes Risk (MBR) using sentence-embedding similarity, and then fine-tunes the model with ORPO on the highest-scoring output as the chosen response and a low-scoring output as the rejected response. Experiments on Llama-3.1 and Qwen-2.5 models across five LongBench tasks report an average SubEM improvement of 4.2 points for Llama-3.1-8B-Instruct (50.8 to 55.0), with Qwen-2.5-14B-Instruct exceeding its 32B variant. The authors also compare SEALONG with prior human- or expert-annotated datasets and find that it outperforms GPT-4o-MuSiQue (55.0 vs. 50.9). Additional analyses examine scoring methods, the number of synthetic examples, the number of samples per question, and short-context performance stability.
Significance. If the reported results are robust, the contribution is significant: SEALONG would demonstrate that consensus-based MBR selection over a model's own sampled outputs can provide effective training supervision for long-context reasoning without human or expert-model annotations. This would lower the cost of data synthesis and offer a practical route to self-improvement. The paper also includes useful control analyses: a token-count check showing that gains are not explained by longer outputs (Table 4), a comparison against several prior datasets with a fair 2K-example budget (Table 5), and an evaluation showing minimal short-context degradation (Table 8). However, the central claim is currently only partially supported because of three load-bearing gaps: MuSiQue is used both as the training-data source and as an evaluation task, no control exists for random selection of self-generated training targets, and no error bars or significance tests are reported. These issues are addressable within the manuscript's scope.
major comments (4)
- [§4.1, Table 2] MuSiQue serves both as the source of training questions (§4.1: 'we leverage the training dataset of MuSiQue') and as one of the five evaluation tasks in Table 2. For Llama-3.1-8B-Instruct, the largest gain is on MuSiQue (49.5 to 58.5, +9.0 points), and the headline 50.8-to-55.0 average includes this in-domain result. The caption of Table 2 correctly frames the other tasks as demonstrating generalization, but the reported average does not. To support the claim of self-improvement in long-context reasoning, report the average over the four out-of-domain tasks separately, and ideally add a held-out multi-hop QA task that does not overlap with MuSiQue in question style or distribution.
- [§3.2, §4.4, Table 7] There is no control where the training target is selected randomly (or by lowest MBR score) from the same N sampled outputs. Table 7 shows that MBR decoding selects better outputs than random at inference time, but that does not establish that MBR-selected training targets are superior to randomly selected ones after fine-tuning. The 4.2-point improvement could therefore arise from fine-tuning on self-generated in-domain outputs alone rather than from the MBR filter. Add an ablation that fine-tunes on randomly sampled outputs (and, ideally, on lowest-scoring outputs) from the same pool, with identical ORPO hyperparameters; if the random control achieves a similar gain, the central claim about MBR-based self-supervision is not supported.
- [Tables 2, 5, 7, 8; Figs. 3–4] All results are reported as point estimates without error bars, multiple seeds, or significance tests. The evaluation sets contain only 150–200 questions per task (Table 3), so a 4.2-point average difference may be within run-to-run or sampling noise; this is especially relevant for the comparison of Qwen-2.5-14B+SEALONG (54.7) versus Qwen-2.5-32B (53.1). Report bootstrap confidence intervals for the main averages, or per-task significance tests (e.g., McNemar's test), and run at least two fine-tuning seeds for the headline results.
- [§3.1, Fig. 1, Limitations] The central premise that 'correct reasoning trajectories typically exhibit higher semantic consistency' is only partially supported. Figure 1 shows that MBR decoding remains far below the oracle sample even at N=128, and the Limitations section acknowledges this gap. This weakness is not fatal by itself, but it makes the missing random-selection control (second major comment) more acute. A quantitative analysis on a small labeled subset of the training questions, reporting how often the MBR-selected output is correct, would strengthen the premise and help interpret the fine-tuning gains.
minor comments (6)
- [§3.1, §4.1] The description of the embedding model is inconsistent: §3.1 says 'a lightweight RoBERTa-based model' while §4.1 says 'jina-embeddings-v3 serving as the sentence embedding model.' Clarify which model is actually used and whether it is part of the self-supervision pipeline.
- [Limitations, §4.3, §4.4] There are several typographical errors: 'Minimum Bayesian Risk' in the Limitations section should read 'Minimum Bayes Risk'; 'dose not' should read 'does not'; 'particularlly' should read 'particularly'; and 'see also in Tab. 1' in §4.4 likely refers to Fig. 1.
- [Table 7] The caption of Table 7 reports the performance of the highest-scoring output for each method but does not include the oracle upper bound. Adding the oracle score would help readers gauge the remaining headroom for the scoring methods.
- [§4.1] The construction of synthetic contexts is underspecified: 'we randomly sample some unrelated documents' — please state how unrelated documents are selected, whether they are drawn from the same MuSiQue corpus, and whether answer leakage across questions is possible.
- [Figs. 3 and 4] The captions for Figures 3 and 4 do not specify the exact metric plotted. State that the vertical axis is the average SubEM over the five LongBench tasks (or otherwise specify the metric).
- [§5] The sentence 'SEALONG first reveals the underestimated potential of LLMs in long-context reasoning' overstates the contribution; consider softening to 'provides evidence for the underestimated potential.'
Circularity Check
No significant circularity: the self-supervision is self-training, not a definitional reduction, and the evaluation is external with no fitted parameter renamed as prediction.
full rationale
SEALONG's chain is empirical: sample outputs from the base LLM, rank them by MBR consistency computed from those same samples, fine-tune with SFT or ORPO, then measure SubEM on held-out LongBench tasks and Open LLM Leaderboard. The training selection score (sentence-embedding similarity to other sampled outputs) is not the evaluation quantity (SubEM against gold answers), so the reported gains are not forced by construction. The self-referential aspect—the model generates its own training targets—is self-training, which can fail if consensus tracks confident errors; the Limitations section explicitly concedes that 'a substantial performance gap remains between the highest MBR-scored output and the oracle sample,' making the premise empirically testable rather than definitional. Citations to prior consensus-based and MBR work are background support, not author-specific uniqueness theorems, and none of the central choices (plan-and-solve prompting, jina-embeddings-v3, ORPO) is justified solely by self-citation. The absence of a random-selection training control is a potential experimental confound, but that is a validity concern, not a circular reduction.
Assumptions & free parameters
free parameters (2)
- Number of samples per question N =
32
- Number of synthetic training examples =
2048
assumptions (5)
- domain assumption Consensus correlates with correctness in sampled LLM outputs
- domain assumption Substring exact match is an adequate evaluation of reasoning quality
- domain assumption Plan-and-solve prompting elicits the best available reasoning from base models
- domain assumption Training on MuSiQue question style transfers to other long-context QA tasks
- domain assumption Sentence-embedding similarity approximates semantic equivalence of reasoning trajectories
Cite this review
Pith. "Pith review of Large Language Models Can Self-Improve in Long-context Reasoning." pith.science (2026). https://pith.science/paper/MTZRIYMU
@misc{pith2026241108147,
author = {Pith},
title = {Pith review of: Large Language Models Can Self-Improve in Long-context Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/MTZRIYMU}},
note = {Machine review of arXiv:2411.08147}
}
abstract
Large language models (LLMs) have achieved substantial progress in processing long contexts but still struggle with long-context reasoning. Existing approaches typically involve fine-tuning LLMs with synthetic data, which depends on annotations from human experts or advanced models like GPT-4, thus restricting further advancements. To address this issue, we investigate the potential for LLMs to self-improve in long-context reasoning and propose \ours, an approach specifically designed for this purpose. This approach is straightforward: we sample multiple outputs for each question, score them with Minimum Bayes Risk, and then apply supervised fine-tuning or preference optimization based on these outputs. Extensive experiments on several leading LLMs demonstrate the effectiveness of \ours, with an absolute improvement of $4.2$ points for Llama-3.1-8B-Instruct. Furthermore, \ours achieves superior performance compared to prior approaches that depend on data produced by human experts or advanced models. We anticipate that this work will open new avenues for self-improvement techniques in long-context scenarios, which are essential for the continual advancement of LLMs.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 2 Pith papers
-
UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp
A unified KG-plus-live-web data pipeline covering all three multimodal BrowseComp information-flow patterns, plus an exploration-degree filter, yields a 35B agent at 54.4 avg accuracy.
-
Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning
MMSci, a new scientific table benchmark and training set, shows that 52K domain-specific table images outperform 150K general-domain images for multimodal numerical reasoning.
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[2]
Chenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu. 2024 a . L -eval: Instituting standardized evaluation for long context language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand. Association for Computational...
2024
-
[3]
Chenxin An, Fei Huang, Jun Zhang, Shansan Gong, Xipeng Qiu, Chang Zhou, and Lingpeng Kong. 2024 b . Training-free long-context scaling of large language models. In Forty-first International Conference on Machine Learning
2024
-
[4]
Chenxin An, Jun Zhang, Ming Zhong, Lei Li, Shansan Gong, Yao Luo, Jingjing Xu, and Lingpeng Kong. 2024 c . Why does the effective context length of llms fall short? arXiv preprint arXiv:2410.18745
arXiv 2024
-
[5]
Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, and Jian-Guang Lou. 2024 d . Make your llm fully utilize the context. arXiv preprint arXiv:2404.16811
arXiv 2024
-
[6]
Anthropic. 2023. https://www.anthropic.com/index/claude-2-1 Anthropic: Introducing claude 2.1
2023
-
[7]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self- RAG : Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations
2024
-
[8]
Yushi Bai, Xin Lv, Jiajie Zhang, Yuze He, Ji Qi, Lei Hou, Jie Tang, Yuxiao Dong, and Juanzi Li. 2024. Longalign: A recipe for long context alignment of large language models. arXiv preprint arXiv:2401.18058
arXiv 2024
Show all 112 references
-
[9]
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. 2023. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508
2023 arXiv
-
[10]
Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. 2023. Open llm leaderboard
2023
-
[11]
Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew Gormley. 2024. Unlimiformer: Long-range transformers with unlimited length input. Advances in Neural Information Processing Systems, 36
2024
-
[12]
Amanda Bertsch, Alex Xie, Graham Neubig, and Matthew R Gormley. 2023. It’s mbr all the way down: Modern generation techniques through the lens of minimum bayes risk. In Proceedings of the Big Picture Workshop, pages 108--122
2023
-
[13]
Bickel and K.A
P.J. Bickel and K.A. Doksum. 1977. Mathematical Statistics: Basic Ideas and Selected Topics. Prentice Hall
1977
-
[14]
S \'e bastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712
2023 arXiv
-
[15]
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595
2023 arXiv
-
[16]
Xinyun Chen, Maxwell Lin, Nathanael Sch \"a rli, and Denny Zhou. 2024 a . Teaching large language models to self-debug. In The Twelfth International Conference on Learning Representations
2024
-
[17]
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2024 b . Longlo RA : Efficient fine-tuning of long-context large language models. In The Twelfth International Conference on Learning Representations
2024
-
[18]
Zhi Chen, Qiguang Chen, Libo Qin, Qipeng Guo, Haijun Lv, Yicheng Zou, Wanxiang Che, Hang Yan, Kai Chen, and Dahua Lin. 2024 c . What are the essential factors in crafting effective long context multi-hop instruction datasets? insights and best practices. arXiv preprint arXiv:2...
2024 arXiv
-
[19]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457
2018 arXiv
-
[20]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
2021 arXiv
-
[21]
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. Ultrafeedback: Boosting language models with high-quality feedback. arXiv preprint arXiv:2310.01377
2023 arXiv
-
[22]
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational ...
2021
-
[23]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2024. Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems, 36
2024
-
[24]
Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, Nanning Zheng, and Furu Wei. 2023. Longnet: Scaling transformers to 1,000,000,000 tokens. arXiv preprint arXiv:2307.02486
2023 arXiv
-
[25]
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. 2024. Longro PE : Extending LLM context window beyond 2 million tokens. In Forty-first International Conference on Machine Learning
2024
-
[26]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[27]
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024. Detecting hallucinations in large language models using semantic entropy. Nature, 630(8017):625--630
2024
-
[28]
Mara Finkelstein and Markus Freitag. 2024. MBR and QE finetuning: Training-time distillation of the best and most expensive decoding methods. In The Twelfth International Conference on Learning Representations
2024
-
[29]
Yao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Yoon Kim, and Hao Peng. 2024. Data engineering for scaling language models to 128k context. arXiv preprint arXiv:2402.10171
2024 arXiv
-
[30]
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I Liao, Kamil \.e Luko s i \=u t \.e , Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al. 2023. The capacity for moral self-correction in large language models. arXiv preprint arXiv:2302.07459
2023 arXiv
-
[31]
Tianyu Gao, Alexander Wettig, Howard Yen, and Danqi Chen. 2024. How to train long-context language models (effectively). arXiv preprint arXiv:2410.02660
2024
-
[32]
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, et al. 2024. Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793
2024 arXiv
-
[33]
Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2024. CRITIC : Large language models can self-correct with tool-interactive critiquing. In The Twelfth International Conference on Learning Representations
2024
-
[34]
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al. 2023. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998
2023 arXiv
-
[35]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations
2021
-
[36]
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6609--6625
2020
-
[37]
Jiwoo Hong, Noah Lee, and James Thorne. 2024. Orpo: Monolithic preference optimization without reference model. arXiv preprint arXiv:2403.07691, 2(4):5
2024 arXiv
-
[38]
Arian Hosseini, Xingdi Yuan, Nikolay Malkin, Aaron Courville, Alessandro Sordoni, and Rishabh Agarwal. 2024. V- ST ar: Training verifiers for self-taught reasoners. In First Conference on Language Modeling
2024
-
[39]
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg. 2024. Ruler: What's the real context size of your long-context language models? arXiv preprint arXiv:2404.06654
2024 arXiv
-
[40]
Jiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2023. Large language models can self-improve. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 1051--1068
2023
-
[41]
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2024. Large language models cannot self-correct reasoning yet. In The Twelfth International Conference on Learning Representations
2024
-
[42]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276
2024 arXiv
-
[43]
Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A Smith, Iz Beltagy, et al. 2023. Camels in a changing climate: Enhancing lm adaptation with tulu 2. arXiv preprint arXiv:2311.10702
2023 arXiv
-
[44]
Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Shuaiwen Leon Song, Samyam Rajbhandari, and Yuxiong He. 2023. Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models. arXiv preprint arXiv:2309.14509
2023 arXiv
-
[45]
Dongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, and Daniel Khashabi. 2024. Self-[in] correct: Llms struggle with refining self-generated responses. arXiv preprint arXiv:2404.04298
2024 arXiv
-
[46]
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024. SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations
2024
-
[47]
Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Zirui Liu, Chia-Yuan Chang, Huiyuan Chen, and Xia Hu. 2024. LLM maybe long LM : Selfextend LLM context window without tuning. In Forty-first International Conference on Machine Learning
2024
-
[48]
Greg Kamradt. 2023. Needle in a haystack - pressure testing llms
2023
-
[49]
Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2024. One thousand and one pairs: A" novel" challenge for long-context language models. arXiv preprint arXiv:2406.16264
2024 arXiv
-
[50]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199--22213
2022
-
[51]
Shankar Kumar and Bill Byrne. 2004. Minimum bayes-risk decoding for statistical machine translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 169--176
2004
-
[52]
Tian Lan, Wenwei Zhang, Chengqi Lyu, Shuaibin Li, Chen Xu, Heyan Huang, Dahua Lin, Xian-Ling Mao, and Kai Chen. 2024 a . Training language models to critique with multi-agent feedback. arXiv preprint arXiv:2410.15287
2024 arXiv
-
[53]
Tian Lan, Wenwei Zhang, Chen Xu, Heyan Huang, Dahua Lin, Kai Chen, and Xian-ling Mao. 2024 b . Criticbench: Evaluating large language models as critic. arXiv preprint arXiv:2402.13764
2024 arXiv
-
[54]
Mosh Levy, Alon Jacoby, and Yoav Goldberg. 2024. Same task, more tokens: the impact of input length on the reasoning performance of large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangko...
2024
-
[55]
Huayang Li, Pat Verga, Priyanka Sen, Bowen Yang, Vijay Viswanathan, Patrick Lewis, Taro Watanabe, and Yixuan Su. 2024 a . A retrieve-then-reason framework for long-context question answering. arXiv preprint arXiv:2410.03227
2024 arXiv
-
[56]
Mo Li, Songyang Zhang, Yunxin Liu, and Kai Chen. 2024 b . Needlebench: Can llms do retrieval and reasoning in 1 million context window? arXiv preprint arXiv:2407.11963
2024
-
[57]
Yanyang Li, Shuo Liang, Michael Lyu, and Liwei Wang. 2024 c . Making long-context language models better multi-hop reasoners. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2462--2475
2024
-
[58]
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al. 2024. Jamba: A hybrid transformer-mamba language model. arXiv preprint arXiv:2403.19887
2024 arXiv
-
[59]
Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74--81
2004
-
[60]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214--3252
2022
-
[61]
Zicheng Lin, Zhibin Gou, Tian Liang, Ruilin Luo, Haowei Liu, and Yujiu Yang. 2024. C ritic B ench: Benchmarking LLM s for critique-correct reasoning. In Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand. Association for Computational Linguistics
2024
-
[62]
Yinhan Liu. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 364
2019 arXiv
-
[63]
Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. 2024. Agentboard: An analytical evaluation board of multi-turn llm agents. arXiv preprint arXiv:2401.13178
2024 arXiv
-
[64]
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2024. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36
2024
-
[65]
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Compu...
2023
-
[66]
Potsawee Manakul, Adian Liusie, and Mark Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9004--9017
2023
-
[67]
OpenAI. 2022. Chatgpt blog post. https://openai.com/blog/chatgpt
2022
-
[68]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 3...
2022
-
[69]
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. 2024. Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies. Transactions of the Association for Computational Linguistics, 1...
2024
-
[70]
Richard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho, Sainbayar Sukhbaatar, and Jason E Weston. 2024. Iterative reasoning preference optimization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[71]
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2024. Ya RN : Efficient context window extension of large language models. In The Twelfth International Conference on Learning Representations
2024
-
[72]
Archiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu, Maryam Fazel-Zarandi, Mohit Bansal, Sainbayar Sukhbaatar, Jason Weston, and Jane Yu. 2024. https://arxiv.org/abs/2411.04109 Self-consistency preference optimization . Preprint, arXiv:2411.04109
2024 arXiv
-
[73]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36
2024
-
[74]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99--106
2021
-
[75]
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36
2024
-
[76]
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael G \"u nther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, et al. 2024. jina-embeddings-v3: Multilingual embeddings with task lora. arXiv preprint arXiv:2409.10173
2024 arXiv
-
[77]
Yutao Sun, Li Dong, Yi Zhu, Shaohan Huang, Wenhui Wang, Shuming Ma, Quanlu Zhang, Jianyong Wang, and Furu Wei. 2024. You only cache once: Decoder-decoder architectures for language models. arXiv preprint arXiv:2405.05254
2024 arXiv
-
[78]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Musique: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539--554
2022
-
[79]
Roy Tromble, Shankar Kumar, Franz Josef Och, and Wolfgang Macherey. 2008. Lattice minimum bayes-risk decoding for statistical machine translation. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, pages 620--629
2008
-
[80]
Kiran Vodrahalli, Santiago Ontanon, Nilesh Tripuraneni, Kelvin Xu, Sanil Jain, Rakesh Shivanna, Jeffrey Hui, Nishanth Dikkala, Mehran Kazemi, Bahare Fatemi, et al. 2024. Michelangelo: Long context evaluations beyond haystacks via latent structure queries. arXiv preprint arXiv:...
2024 arXiv
-
[81]
Jun Wang, Eleftheria Briakou, Hamid Dadkhahi, Rishabh Agarwal, Colin Cherry, and Trevor Cohn. 2024 a . Don't throw away data: Better sequence knowledge distillation. arXiv preprint arXiv:2407.10456
2024 arXiv
-
[82]
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023 a . Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational ...
2023
-
[83]
Minzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, et al. 2024 b . Leave no document behind: Benchmarking long-context llms with extended multi-doc qa. arXiv preprint arXiv:2406.17419
2024 arXiv
-
[84]
Tianduo Wang, Shichen Li, and Wei Lu. 2024 c . Self-training with direct preference optimization improves chain-of-thought reasoning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11917--11928
2024
-
[85]
Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei. 2024 d . Augmenting language models with long-term memory. Advances in Neural Information Processing Systems, 36
2024
-
[86]
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 b . Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Lin...
2023
-
[87]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[88]
Ian Wu, Patrick Fernandes, Amanda Bertsch, Seungone Kim, Sina Pakazad, and Graham Neubig. 2024. Better instruction-following through minimum bayes risk. arXiv preprint arXiv:2410.02902
2024 arXiv
-
[89]
Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, and Christian Szegedy. 2022. Memorizing transformers. In International Conference on Learning Representations
2022
-
[90]
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao, Min-Yen Kan, Junxian He, and Michael Xie. 2024. Self-evaluation guided beam search for reasoning. Advances in Neural Information Processing Systems, 36
2024
-
[91]
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, S...
2024
-
[92]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng...
2024 arXiv
-
[93]
Guangyu Yang, Jinghong Chen, Weizhe Lin, and Bill Byrne. 2024 b . Direct preference optimization for neural machine translation with minimum bayes risk decoding. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistic...
2024
-
[94]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language ...
2018
-
[95]
Tianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun, Yi Zhu, Gao Huang, and Furu Wei. 2024. Differential transformer. arXiv preprint arXiv:2410.05258
2024 arXiv
-
[96]
Howard Yen, Tianyu Gao, and Danqi Chen. 2024 a . Long-context language modeling with parallel context encoding. arXiv preprint arXiv:2402.16617
2024 arXiv
-
[97]
Howard Yen, Tianyu Gao, Minmin Hou, Ke Ding, Daniel Fleischer, Peter Izasak, Moshe Wasserblat, and Danqi Chen. 2024 b . Helmet: How to evaluate long-context language models effectively and thoroughly. arXiv preprint arXiv:2410.02694
2024 arXiv
-
[98]
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason E Weston. 2024. Self-rewarding language models. In Forty-first International Conference on Machine Learning
2024
-
[99]
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. ST ar: Bootstrapping reasoning with reasoning. In Advances in Neural Information Processing Systems
2022
-
[100]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4800
2019
-
[101]
Dan Zhang, Sining Zhoubian, Yisong Yue, Yuxiao Dong, and Jie Tang. 2024 a . Rest-mcts*: Llm self-training via process reward guided tree search. arXiv preprint arXiv:2406.03816
2024 arXiv
-
[102]
Jiajie Zhang, Zhongni Hou, Xin Lv, Shulin Cao, Zhenyu Hou, Yilin Niu, Lei Hou, Yuxiao Dong, Ling Feng, and Juanzi Li. 2024 b . Longreward: Improving long-context large language models with ai feedback. arXiv preprint arXiv:2410.21252
2024 arXiv
-
[103]
Peitian Zhang, Ninglu Shao, Zheng Liu, Shitao Xiao, Hongjin Qian, Qiwei Ye, and Zhicheng Dou. 2024 c . Extending llama-3's context ten-fold overnight. arXiv preprint arXiv:2404.19553
2024 arXiv
-
[104]
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675
2019 arXiv
-
[105]
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, and Maosong Sun. 2024 d . B ench: Extending long context evaluation beyond 100 K tokens. In Proceedings of the 62nd Annual Meeting of the Association for Compu...
2024
-
[106]
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024 a . Wildchat: 1m chat GPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations
2024
-
[107]
Xinran Zhao, Hongming Zhang, Xiaoman Pan, Wenlin Yao, Dong Yu, Tongshuang Wu, and Jianshu Chen. 2024 b . Fact-and-reflection ( F a R ) improves confidence calibration of large language models. In Findings of the Association for Computational Linguistics ACL 2024, Bangkok, Thai...
2024
-
[108]
Tianyang Zhong, Zhengliang Liu, Yi Pan, Yutong Zhang, Yifan Zhou, Shizhe Liang, Zihao Wu, Yanjun Lyu, Peng Shu, Xiaowei Yu, et al. 2024. Evaluation of openai o1: Opportunities and challenges of agi. arXiv preprint arXiv:2409.18486
2024
-
[109]
Denny Zhou, Nathanael Sch \"a rli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi. 2023. Least-to-most prompting enables complex reasoning in large language models. In The Eleventh International Conference...
2023
-
[110]
Dawei Zhu, Nan Yang, Liang Wang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. 2024. Po SE : Efficient context window extension of LLM s via positional skip-wise training. In The Twelfth International Conference on Learning Representations
2024
-
[111]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[112]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.