REVIEW 4 major objections 5 minor 1 cited by
Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that prefix-tuning a 1.5B math reasoning model for diversity reaches 80% accuracy with 32 samples rather than 256.
desk verdict A competent survey plus a small, honest empirical study whose central diversity claim is not yet supported by the evidence, and which contains a clear internal inconsistency in the reported numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is prefix fine-tuning over a curated data mixture. The training set is 90% responses from Qwen2.5-Math-1.5B generated with a prompt that asks for the initial step toward solving the question, which is meant to encourage varied reasoning beginnings, plus 10% of the target model's own outputs to prevent forgetting. All examples are truncated to 512 tokens and only prefix parameters are updated while the rest of the model stays frozen; at inference, the tuned model is evaluated with best-of-N sampling and majority voting.
What would settle it
Compute self-BLEU and pairwise entropy over the N=32 candidate solutions from ADAPT and from DeepSeek-R1-Distill-Qwen-1.5B at temperature 0.8; if ADAPT's accuracy advantage appears while its diversity metrics are no higher, the paper's mechanism is wrong. A control trained on 90% Qwen responses without the initial-step prompt that still gains accuracy would likewise show that transfer, not diversity, drives the result.
Extended reading notes
Core claim
On its own terms, the paper shows that a small prefix fine-tune can restore the output variety that reasoning distillation removes, and that this variety is what makes sampling-based test-time scaling pay off. With only the first 512 tokens of each training example and a 90/10 mixture of Qwen2.5-Math-1.5B and DeepSeek-R1-Distill-Qwen-1.5B outputs, ADAPT reaches 80% accuracy at N=32, versus N=256 for the unmodified distilled model, and peaks at 81.0%. The paper takes this as evidence that generative diversity is a key enabler of test-time scaling, not just a side effect of sampling temperature.
Load-bearing premise
The paper assumes its 90/10 training mixture increases the variety of solution paths the target model produces, and that the accuracy gains come from that variety rather than from borrowed reasoning ability, yet it never measures diversity directly.
Editorial extensions
If this is right
- Reasoning distillation, while raising single-sample accuracy, can quietly suppress the output diversity that sampling-based test-time scaling depends on.
- ADAPT reaches 80% accuracy at N=32 rather than N=256, an 8x reduction in the inference budget needed to hit the same threshold.
- Most of ADAPT's gains arrive before N=32, so small compute budgets capture most of the benefit of test-time scaling.
- At N=256, ADAPT's peak accuracy of 81.0% slightly exceeds the distilled baseline's 80.8%, so the efficiency gain is not bought with peak performance.
Reading between the lines
- If diversity is the active ingredient, the same prefix-tuning recipe should transfer to other low-diversity reasoning models and other aggregation rules such as verifier reranking; the paper only tests majority voting on math.
- A cleaner test would compare ADAPT against a control trained on 90% Qwen responses with the standard chat template, isolating the contribution of the initial-step prompt from raw knowledge transfer.
- Because only prefix parameters are updated, the method should be cheap enough to apply at larger scale, but the paper does not test this directly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper combines a survey of test-time scaling (TTS) methods—taxonomized into sampling, search, and trajectory optimization—with a proposed method, ADAPT, that applies prefix tuning to a 1.5B distilled reasoning model using a 90/10 mixture of Qwen2.5-Math-1.5B and DeepSeek-R1-Distill-Qwen-1.5B outputs. The empirical claim is that ADAPT reaches 80% accuracy with N=32 best-of-N majority-vote samples, an 8x efficiency gain over the DeepSeek baseline that requires N=256, and that this gain stems from increased generative diversity. The paper also includes future-directions discussion and an appendix with training details.
Significance. If the central claim holds, the paper would provide a simple and parameter-efficient way to improve test-time scaling in small reasoning-optimized models, with potentially broad applicability. The survey component offers a useful, if not deeply novel, organization of recent TTS literature. The authors are explicitly honest about limitations, and the reported efficiency numbers, if reproducible, are practically interesting. However, the key attribution of the gain to diversity is not directly supported by any measurement, and the experimental evidence rests on a single run on an unnamed benchmark. Consequently, the significance is conditional on additional validation that is within the scope of a revision.
major comments (4)
- [Section 5.4, Table 1] The text states that 'At N=16, it already surpasses the distilled baseline at N=256,' but Table 1 reports ADAPT at 78.2% and DeepSeek Qwen-1.5B at 80.8% at those respective sample counts. The claim is contradicted by the paper's own table and must be corrected or substantiated with the actual accuracy at N=16.
- [Section 7 (Limitations), Section 4.1] The paper acknowledges that it does not directly measure diversity (e.g., self-BLEU, pairwise entropy) and that the diversity effect is 'inferred only through indirect accuracy gains.' This is load-bearing because the title, method name, and Section 5.5 all attribute ADAPT's efficiency improvement to increased diversity, yet the training mixture is 90% Qwen2.5 responses, making content transfer or answer-format transfer a live confound. Without any diversity metric computed on generated outputs, the accuracy curves in Table 1 cannot distinguish the diversity hypothesis from a content-transfer hypothesis. The revision should measure diversity directly (e.g., self-BLEU, n-gram overlap, or pairwise entropy of sampled solutions at a given N) or explicitly soften the causal claim.
- [Section 5.1, Experiments] The evaluation benchmark is described only as 'akin to MATH-500' without naming the exact dataset, and there is no mention of random seeds, multiple runs, or variance estimates. For a paper whose central claim is a quantitative efficiency ratio (8x) derived from a single accuracy threshold, this is insufficient experimental reporting. The authors should specify the exact benchmark, provide results over multiple seeds with standard errors or confidence intervals, and report the actual accuracy at N=32 (the text infers Min N but does not show the measured value).
- [Section 4.1, Section 5.4] The 90/10 mixture ratio and the custom prompt template are free parameters of the method, yet no ablations are performed to isolate the contribution of each component. Without ablations (e.g., 100% DeepSeek data with the same prefix tuning, or 100% Qwen data without the prompt template), it is not possible to tell whether the observed gains are due to diversity-oriented data selection, the prompt template, or simply the added Qwen-derived training signal. Reporting such ablations would also help test the diversity mechanism.
minor comments (5)
- [Abstract, Section 5.4] The abstract says 'eight times less compute' while the body says '8× improvement' in sample count; please use consistent wording and clarify that the comparison is in number of samples, not wall-clock compute.
- [Figure 5 caption] The caption reads 'Marginal gain per generation for across models' and should be corrected to '...for all models'.
- [Section A.5.2] The description of the dataset split ('combined dataset is shuffled and split into 90% for training and 10% for testing') is ambiguous relative to Section 4.1's 90/10 mixture of model outputs; clarify whether the 90/10 refers to data source mixture or train/test split.
- [Appendix A.1, Figure 6] Several entries in the methodology structure are duplicated (e.g., Coconut, Latent-Thought, CODI, and Looped transformer appear under both 'Search/Hidden layer search' and 'Search/Self-improvement'); these should be listed under the most appropriate category only.
- [References] Some references are incomplete or nonstandard: 'Beeching et al.' appears without a year or full citation, 'Chen et al.' in the Safety paragraph of Section 6 lacks a reference, and several entries contain 'and 1 others' instead of the full author list. Please review and complete the bibliography.
Circularity Check
No circularity in the efficiency measurement; the diversity attribution is a partly self-referential interpretation, acknowledged in the Limitations, so only a low score.
-
other
[Section 5.5 and Limitations (p. 8-9); also Sections 4.1 and 5.3]
"Although ADAPT improves sample efficiency, it does not directly optimize diversity metrics (e.g., self-BLEU, pairwise entropy), and its diversity-enhancing effect is inferred only through indirect accuracy gains. ... Our results validate the central hypothesis: enhancing output diversity improves both the accuracy and efficiency of reasoning models under Best-of-N sampling."
The paper's central interpretive claim is that ADAPT's gains validate the diversity hypothesis, but ADAPT's diversity enhancement is never measured. The training mixture is labeled 'diverse' because Qwen2.5-Math-1.5B is assumed to have higher generative diversity, and that assumption is supported in Section 5.3 by the same TTS scaling behavior the hypothesis is meant to explain. The accuracy gains are then used to confirm the diversity hypothesis, while the diversity increase is inferred only from those accuracy gains. This closes an interpretive loop: diversity is inferred from the accuracy improvement, and the accuracy improvement is explained by diversity. The headline empirical efficiency comparison is not circular, which is why the score is low.
full rationale
The headline result—ADAPT reaches 80% accuracy at N=32 versus N=256 for DeepSeek-Qwen-1.5B—is an empirical measurement against external baselines under a fixed Best-of-N majority-voting protocol. It is not derived by fitting or by construction, so the central efficiency claim is self-contained. The only circular element is interpretive: the claim that the results validate the diversity hypothesis rests on the unmeasured assumption that the 90% Qwen training mixture increases diversity. The Limitations section explicitly concedes that no direct diversity metric is measured and that the diversity effect is 'inferred only through indirect accuracy gains.' This makes the diversity explanation self-referential, but it does not invalidate the measured accuracy comparison. Score 2 reflects a minor interpretive circularity, not a construction-level circularity.
Assumptions & free parameters
free parameters (4)
- data_mixture_ratio =
90% Qwen / 10% DeepSeek
- prefix_truncation_length =
512 tokens
- training_epochs =
3
- learning_rate =
5e-6
assumptions (3)
- domain assumption Majority voting over N samples is a valid test-time scaling strategy and accuracy improves with N.
- domain assumption Qwen2.5-Math-1.5B responses with the custom 'initial step' prompt are a valid source of diverse reasoning prefixes.
- domain assumption The evaluation benchmark is representative of mathematical reasoning despite not being named.
Cite this review
Pith. "Pith review of Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning." pith.science (2026). https://pith.science/paper/KOO3ODKE
@misc{pith2026250604611,
author = {Pith},
title = {Pith review of: Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/KOO3ODKE}},
note = {Machine review of arXiv:2506.04611}
}
read the original abstract
Test-Time Scaling (TTS) improves the reasoning performance of Large Language Models (LLMs) by allocating additional compute during inference. We conduct a structured survey of TTS methods and categorize them into sampling-based, search-based, and trajectory optimization strategies. We observe that reasoning-optimized models often produce less diverse outputs, which limits TTS effectiveness. To address this, we propose ADAPT (A Diversity Aware Prefix fine-Tuning), a lightweight method that applies prefix tuning with a diversity-focused data strategy. Experiments on mathematical reasoning tasks show that ADAPT reaches 80% accuracy using eight times less compute than strong baselines. Our findings highlight the essential role of generative diversity in maximizing TTS effectiveness.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling
Perturbation-based selection does not beat a format-matched control that spends the same short-answer budget on the unperturbed image, so reported gains against CoT-only majority voting are mostly a decoding-format effect.
Reference graph
Works this paper leans on
-
[1]
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. https://arxiv.org/abs/2108.07732 Program synthesis with large language models . Preprint, arXiv:2108.07732
arXiv 2021
-
[2]
https://huggingface.co/spaces/HuggingFaceH4/blogpost-scaling-test-time-compute Scaling test-time compute with open models
Edward Beeching, Lewis Tunstall, and Sasha Rush. https://huggingface.co/spaces/HuggingFaceH4/blogpost-scaling-test-time-compute Scaling test-time compute with open models
-
[3]
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2024. https://doi.org/10.1609/aaai.v38i16.29720 Graph of thoughts: Solving elaborate problems with large language models . Proceedings of the AAAI Conference on Artificial ...
-
[4]
Zhenni Bi, Kai Han, Chuanjian Liu, Yehui Tang, and Yunhe Wang. 2025. https://arxiv.org/abs/2412.09078 Forest-of-thought: Scaling test-time compute for enhancing llm reasoning . Preprint, arXiv:2412.09078
arXiv 2025
-
[5]
Houda Bouamor, Juan Pino, and Kalika Bali, editors. 2023. https://aclanthology.org/2023.emnlp-main.0/ Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Singapore
2023
-
[6]
Feng Chen, Allan Raventos, Nan Cheng, Surya Ganguli, and Shaul Druckmann. 2025 a . https://arxiv.org/abs/2502.07154 Rethinking fine-tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning . Preprint, arXiv:2502.07154
arXiv 2025
-
[7]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, and 39 others. 2021. https://arxiv.org/abs/2107.03374 Evaluating large lang...
arXiv 2021
-
[8]
Xinghao Chen, Zhijing Sun, Wenjin Guo, Miaoran Zhang, Yanjun Chen, Yirong Sun, Hui Su, Yijie Pan, Dietrich Klakow, Wenjie Li, and Xiaoyu Shen. 2025 b . https://arxiv.org/abs/2502.18001 Unveiling the key factors for distilling chain-of-thought reasoning . Preprint, arXiv:2502.18001
arXiv 2025
Show all 106 references
-
[9]
Reasoning models don’t always say what they think
Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner Fabien Roger Vlad Mikulik, Sam Bowman, Jan Leike Jared Kaplan, and 1 others. Reasoning models don’t always say what they think
-
[10]
Jie Cheng, Ruixi Qiao, Lijun Li, Chao Guo, Junle Wang, Gang Xiong, Yisheng Lv, and Fei-Yue Wang. 2025. https://arxiv.org/abs/2504.15275 Stop summation: Min-form credit assignment is all process reward model needs for reasoning . Preprint, arXiv:2504.15275
2025
-
[11]
Yinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang, Bo Dai, Sridhar Thiagarajan, Craig Boutilier, Rishabh Agarwal, Aviral Kumar, and Aleksandra Faust. 2024. https://arxiv.org/abs/2412.15287 Inference-aware fine-tuning for best-of-n sampling in large language models . P...
2024
-
[12]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, and 1 others. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311
2022 arXiv
-
[13]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://arxiv.org/abs/1803.05457 Think you have solved question answering? try arc, the ai2 reasoning challenge . Preprint, arXiv:1803.05457
2018 arXiv
-
[14]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . ...
2021 arXiv
-
[15]
Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Chun Cheung, Jan Kautz, Carlos Guestrin, Tatsunori Hashimoto, Sanmi Koyejo, Yejin Choi, Yu Sun, and Xiaolong Wang. 2025. https://arxiv.org/abs/2504.05298 One-minute video generation wi...
2025 arXiv
-
[16]
Xingyu Dang, Christina Baek, Kaiyue Wen, Zico Kolter, and Aditi Raghunathan. 2025. https://arxiv.org/abs/2504.10478 Weight ensembling improves reasoning in language models . Preprint, arXiv:2504.10478
2025
-
[17]
DeepSeek-AI. 2025. https://arxiv.org/abs/2501.12948 Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning . Preprint, arXiv:2501.12948
2025 arXiv
-
[18]
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. 2023. https://arxiv.org/abs/2305.15408 Towards revealing the mystery behind chain of thought: A theoretical perspective . Preprint, arXiv:2305.15408
2023 arXiv
-
[19]
Sicheng Feng, Gongfan Fang, Xinyin Ma, and Xinchao Wang. 2025. https://arxiv.org/abs/2504.10903 Efficient reasoning models: A survey . Preprint, arXiv:2504.10903
2025
-
[20]
Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang. 2024. https://arxiv.org...
2024 arXiv
-
[21]
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. https://arxiv.org/abs/2211.10435 Pal: Program-aided language models . Preprint, arXiv:2211.10435
2023 arXiv
-
[22]
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. https://doi.org/10.1162/tacl_a_00370 Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies . Transactions of the Association for Computational Lingu...
2021 doi
-
[23]
Anna Goldie, Azalia Mirhoseini, Hao Zhou, Irene Cai, and Christopher D. Manning. 2025. https://arxiv.org/abs/2504.04736 Synthetic data generation & multi-step rl for reasoning & tool use . Preprint, arXiv:2504.04736
2025 arXiv
-
[24]
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. 2024. https://arxiv.org/abs/2412.06769 Training large language models to reason in a continuous latent space . Preprint, arXiv:2412.06769
2024 arXiv
-
[25]
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. 2024. https://arxiv.org/abs/2402.14008 Olympiadbench: A challenging benchmark for promoting agi with olym...
2024 arXiv
-
[26]
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021 a . Measuring coding challenge competence with apps. NeurIPS
2021
-
[27]
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021 b . Aligning ai with shared human values. Proceedings of the International Conference on Learning Representations (ICLR)
2021
-
[28]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 c . Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR)
2021
-
[29]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 d . Measuring mathematical problem solving with the math dataset. NeurIPS
2021
-
[30]
Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, and Matthias Bethge. 2025. https://arxiv.org/abs/2504.07086 A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility . Preprint, arXiv:2504.07086
2025
-
[31]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osin...
2022 arXiv
-
[32]
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023. https://arxiv.org/abs/2305.02301 Distilling step-by-step! outperforming larger language models with less training data and smal...
2023 arXiv
-
[33]
Audrey Huang, Adam Block, Qinghua Liu, Nan Jiang, Akshay Krishnamurthy, and Dylan J. Foster. 2025 a . https://arxiv.org/abs/2503.21878 Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment . Preprint, arXiv:2503.21878
2025 arXiv
-
[34]
Chengsong Huang, Langlin Huang, Jixuan Leng, Jiacheng Liu, and Jiaxin Huang. 2025 b . https://arxiv.org/abs/2503.00031 Efficient test-time scaling via self-calibration . Preprint, arXiv:2503.00031
2025 arXiv
-
[35]
Jianhao Huang, Zixuan Wang, and Jason D. Lee. 2025 c . https://arxiv.org/abs/2502.21212 Transformers learn to implement multi-step gradient descent with chain of thought . Preprint, arXiv:2502.21212
2025 arXiv
-
[36]
Xiaoke Huang, Juncheng Wu, Hui Liu, Xianfeng Tang, and Yuyin Zhou. 2025 d . https://arxiv.org/abs/2504.00869 m1: Unleash the potential of test-time scaling for medical reasoning with large language models . Preprint, arXiv:2504.00869
2025
-
[37]
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024. https://arxiv.org/abs/2403.07974 Livecodebench: Holistic and contamination free evaluation of large language models for code . Preprint, a...
2024 arXiv
-
[38]
Metaxas, and Tong Che
Can Jin, Hongwu Peng, Qixin Zhang, Yujin Tang, Dimitris N. Metaxas, and Tong Che. 2025. https://arxiv.org/abs/2504.09772 Two heads are better than one: Test-time scaling of multi-agent collaborative reasoning . Preprint, arXiv:2504.09772
2025 arXiv
-
[39]
Manuj Kant, Manav Kant, Marzieh Nabi, Preston Carlson, and Megan Ma. 2024. https://arxiv.org/abs/2410.09904 Equitable access to justice: Logical llms show promise . Preprint, arXiv:2410.09904
2024 arXiv
-
[40]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. https://arxiv.org/abs/2001.08361 Scaling laws for neural language models . Preprint, arXiv:2001.08361
2020 arXiv
-
[41]
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. https://doi.org/10.18653/v1/N16-1136 MAWPS : A math word problem repository . In Proceedings of the 2016 Conference of the North A merican Chapter of the Association for Computational L...
2016 doi
-
[42]
Deqian Kong, Minglu Zhao, Dehong Xu, Bo Pang, Shu Wang, Edouardo Honig, Zhangzhang Si, Chuan Li, Jianwen Xie, Sirui Xie, and Ying Nian Wu. 2025. https://arxiv.org/abs/2502.01567 Scalable language models with posterior inference of latent thought vectors . Preprint, arXiv:2502.01567
2025 arXiv
-
[43]
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal Behbahani, and Aleksandra Faust. 2025...
2025
-
[44]
Komal Kumar, Tajamul Ashraf, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, Phillip H. S. Torr, Fahad Shahbaz Khan, and Salman Khan. 2025 b . https://arxiv.org/abs/2502.21321 Llm post-training: A deep dive into reasoning large language mod...
2025 arXiv
-
[45]
Fangyu Lei, Qian Liu, Yiming Huang, Shizhu He, Jun Zhao, and Kang Liu. 2024. https://arxiv.org/abs/2310.15147 S3eval: A synthetic, scalable, systematic evaluation suite for large language models . Preprint, arXiv:2310.15147
2024 arXiv
-
[46]
Xinzhe Li. 2025. https://openreview.net/forum?id=x9VQFjtOPS A survey on LLM test-time compute via search: Tasks, LLM profiling, search algorithms, and relevant frameworks . Transactions on Machine Learning Research
2025
-
[47]
Yanyang Li, Michael Lyu, and Liwei Wang. 2025. Learning to reason from feedback at test-time
2025
-
[48]
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma. 2024. https://arxiv.org/abs/2402.12875 Chain of thought empowers transformers to solve inherently serial problems . Preprint, arXiv:2402.12875
2024 arXiv
-
[49]
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let's verify step by step. arXiv preprint arXiv:2305.20050
2023 arXiv
-
[50]
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. https://openreview.net/forum?id=v8L0pN6EOi Let's verify step by step . In The Twelfth International Conference on Learning Re...
2024
-
[51]
Gonzalez
Kevin Lin, Charlie Snell, Yu Wang, Charles Packer, Sarah Wooders, Ion Stoica, and Joseph E. Gonzalez. 2025. https://arxiv.org/abs/2504.13171 Sleep-time compute: Beyond inference scaling at test-time . Preprint, arXiv:2504.13171
2025 arXiv
-
[52]
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. https://doi.org/10.18653/v1/P17-1015 Program induction by rationale generation: Learning to solve and explain algebraic word problems . In Proceedings of the 55th Annual Meeting of the Association for Computational ...
2017 doi
-
[53]
Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023. https://arxiv.org/abs/2210.10749 Transformers learn shortcuts to automata . Preprint, arXiv:2210.10749
2023 arXiv
-
[54]
Guanlin Liu, Anand Ramachandran, Tanmay Gangwani, Yan Fu, and Abhinav Sethy. 2025. https://arxiv.org/abs/2502.17717 Knowledge distillation with training wheels . Preprint, arXiv:2502.17717
2025 arXiv
-
[55]
Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica. 2025. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl. https://pretty-radio-b75.notion.site/DeepScaleR-Su...
2025
-
[56]
Junyu Ma, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, and Dong Yu. 2025 a . https://arxiv.org/abs/2505.03320 Recall with reasoning: Chain-of-thought distillation for mamba's long-context memory and extrapolation . Preprint, arXiv:2505.03320
2025 arXiv
-
[57]
Xiao Ma, Yuhui Tao, Yuhan Zhang, Zexuan Ji, Yizhe Zhang, and Qiang Chen. 2024. https://arxiv.org/abs/2406.17608 Test-time generative augmentation for medical image segmentation . Preprint, arXiv:2406.17608
2024
-
[58]
Yingwei Ma, Yongbin Li, Yihong Dong, Xue Jiang, Rongyu Cao, Jue Chen, Fei Huang, and Binhua Li. 2025 b . https://arxiv.org/abs/2503.23803 Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute . Preprint, arXiv:2503.23803
2025 arXiv
-
[59]
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, and 1 others. 2023. Self-refine: Iterative refinement with self-feedback
2023
-
[60]
William Merrill and Ashish Sabharwal. 2024. https://arxiv.org/abs/2310.07923 The expressive power of transformers with chain of thought . Preprint, arXiv:2310.07923
2024 arXiv
-
[61]
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. 2024. https://arxiv.org/abs/2410.05229 Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models . Preprint, arXiv:2410.05229
2024 arXiv
-
[62]
NVIDIA, :, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, and 60 others. 202...
2025 arXiv
-
[63]
OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
2023 arXiv
-
[64]
OpenAI . 2024 a . https://openai.com/index/learning-to-reason-with-llms/ Learning to reason with llms . Accessed: 2025-05-18
2024
-
[65]
OpenAI . 2024 b . https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf Openai o3 and o4-mini system card . Technical report, OpenAI. Accessed: 2025-04-22
2024
-
[66]
Qianjun Pan, Wenkai Ji, Yuyang Ding, Junsong Li, Shilian Chen, Junyi Wang, Jie Zhou, Qin Chen, Min Zhang, Yulan Wu, and Liang He. 2025. https://arxiv.org/abs/2505.02665 A survey of slow thinking-based reasoning llms using reinforced learning and inference-time scaling law . Pr...
2025 arXiv
-
[67]
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. https://doi.org/10.18653/v1/2021.naacl-main.168 Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Ling...
2021 doi
-
[68]
Yuxiao Qu, Matthew Y. R. Yang, Amrith Setlur, Lewis Tunstall, Edward Emanuel Beeching, Ruslan Salakhutdinov, and Aviral Kumar. 2025. https://arxiv.org/abs/2503.07572 Optimizing test-time compute via meta reinforcement fine-tuning . Preprint, arXiv:2503.07572
2025 arXiv
-
[69]
Yuxiao Qu, Tianjun Zhang, Naman Garg, and Aviral Kumar. 2024. Recursive introspection: Teaching language model agents how to self-improve
2024
-
[70]
Leonardo Ranaldi, Marco Valentino, Alexander Polonsky, and Andrè Freitas. 2025. https://arxiv.org/abs/2502.12616 Improving chain-of-thought reasoning via quasi-symbolic abstractions . Preprint, arXiv:2502.12616
2025
-
[71]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2024. https://openreview.net/forum?id=Ti67584b98 GPQA : A graduate-level google-proof q&a benchmark . In First Conference on Language Modeling
2024
-
[72]
Abulhair Saparov and He He. 2023. https://openreview.net/forum?id=qFVVBzXxR2V Language models are greedy reasoners: A systematic formal analysis of chain-of-thought . In The Eleventh International Conference on Learning Representations
2023
-
[73]
Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J. Reddi. 2025. https://arxiv.org/abs/2502.17416 Reasoning with latent thoughts: On the power of looped transformers . Preprint, arXiv:2502.17416
2025 arXiv
-
[74]
Amrith Setlur, Nived Rajaraman, Sergey Levine, and Aviral Kumar. 2025. https://arxiv.org/abs/2502.12118 Scaling test-time compute without verification or rl is suboptimal . Preprint, arXiv:2502.12118
2025 arXiv
-
[75]
Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, and Yulan He. 2025. https://arxiv.org/abs/2502.21074 Codi: Compressing chain-of-thought into continuous space via self-distillation . Preprint, arXiv:2502.21074
2025 arXiv
-
[76]
Kumar Shridhar, Alessandro Stolfo, and Mrinmaya Sachan. 2023. https://arxiv.org/abs/2212.00193 Distilling reasoning capabilities into smaller language models . Preprint, arXiv:2212.00193
2023 arXiv
-
[77]
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. https://arxiv.org/abs/2408.03314 Scaling llm test-time compute optimally can be more effective than scaling model parameters . Preprint, arXiv:2408.03314
2024 arXiv
-
[78]
Benedikt Stroebl, Sayash Kapoor, and Arvind Narayanan. 2024. https://arxiv.org/abs/2411.17501 Inference scaling flaws: The limits of llm resampling with imperfect verifiers . Preprint, arXiv:2411.17501
2024
-
[79]
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and Xia Hu. 2025. https://arxiv.org/abs/2503.16419 Stop overthinking: A survey on efficient reasoning for large language models . Preprint, arXiv...
2025 arXiv
-
[80]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/v1/N19-1421 C ommonsense QA : A question answering challenge targeting commonsense knowledge . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associa...
2019 doi
-
[81]
Fengwei Teng, Zhaoyang Yu, Quan Shi, Jiayi Zhang, Chenglin Wu, and Yuyu Luo. 2025. https://arxiv.org/abs/2502.12018 Atom of thoughts for markov llm test-time scaling . Preprint, arXiv:2502.12018
2025
-
[82]
Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yunjie Ji, Yiping Peng, Han Zhao, and Xiangang Li. 2025. https://arxiv.org/abs/2503.19855 Think twice: Enhancing llm reasoning by scaling multi-round test-time thinking . Preprint, arXiv:2503.19855
2025 arXiv
-
[83]
Guiyao Tie, Zeli Zhao, Dingjie Song, Fuyang Wei, Rong Zhou, Yurou Dai, Wen Yin, Zhejian Yang, Jiangyue Yan, Yao Su, Zhenhan Dai, Yifeng Xie, Yihan Cao, Lichao Sun, Pan Zhou, Lifang He, Hechang Chen, Yu Zhang, Qingsong Wen, and 7 others. 2025. https://arxiv.org/abs/2503.06072 A...
2025 arXiv
-
[84]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, and 1 others. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971
2023 arXiv
-
[85]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30
2017
-
[86]
Jiankang Wang, Jianjun Xu, Xiaorui Wang, Yuxin Wang, Mengting Xing, Shancheng Fang, Zhineng Chen, Hongtao Xie, and Yongdong Zhang. 2025 a . https://arxiv.org/abs/2412.08864 A graph-based synthetic data pipeline for scaling high-quality reasoning instructions . Preprint, arXiv:...
2025
-
[87]
Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou. 2024 a . https://arxiv.org/abs/2406.04692 Mixture-of-agents enhances large language model capabilities . Preprint, arXiv:2406.04692
2024 arXiv
-
[88]
Rui Wang, Hongru Wang, Boyang Xue, Jianhui Pang, Shudong Liu, Yi Chen, Jiahao Qiu, Derek Fai Wong, Heng Ji, and Kam-Fai Wong. 2025 b . https://arxiv.org/abs/2503.24377 Harnessing the reasoning economy: A survey of efficient reasoning for large language models . Preprint, arXiv...
2025 arXiv
-
[89]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. https://arxiv.org/abs/2203.11171 Self-consistency improves chain of thought reasoning in language models . Preprint, arXiv:2203.11171
2023 arXiv
-
[90]
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. 2024 b . https://arxiv.org/abs/2406.01574 Mmlu-pro: A more robust ...
2024 arXiv
-
[91]
Sean Welleck, Amanda Bertsch, Matthew Finlayson, Hailey Schoelkopf, Alex Xie, Graham Neubig, Ilia Kulikov, and Zaid Harchaoui. 2024. https://arxiv.org/abs/2406.16838 From decoding to meta-generation: Inference-time algorithms for large language models . Preprint, arXiv:2406.16838
2024 arXiv
-
[92]
Lillicrap, Kenji Kawaguchi, and Michael Shieh
Yuxi Xie, Anirudh Goyal, Wenyue Zheng, Min-Yen Kan, Timothy P. Lillicrap, Kenji Kawaguchi, and Michael Shieh. 2024. https://arxiv.org/abs/2405.00451 Monte carlo tree search boosts reasoning via iterative preference learning . Preprint, arXiv:2405.00451
2024 arXiv
-
[93]
Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, Chenyang Shao, Yuwei Yan, Qinglong Yang, Yiwen Song, Sijian Ren, Xinyuan Hu, Yu Li, Jie Feng, Chen Gao, and Yong Li. 2025 a . https://arxiv.or...
2025 arXiv
-
[94]
Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He. 2025 b . https://arxiv.org/abs/2502.18600 Chain of draft: Thinking faster by writing less . Preprint, arXiv:2502.18600
2025 arXiv
-
[95]
Wenkai Yang, Shuming Ma, Yankai Lin, and Furu Wei. 2025. https://arxiv.org/abs/2502.18080 Towards thinking-optimal scaling of test-time compute for llm reasoning . Preprint, arXiv:2502.18080
2025
-
[96]
Griffiths, Yuan Cao, and Karthik Narasimhan
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. https://arxiv.org/abs/2305.10601 Tree of thoughts: Deliberate problem solving with large language models . Preprint, arXiv:2305.10601
2023 arXiv
-
[97]
Huifeng Yin, Yu Zhao, Minghao Wu, Xuanfan Ni, Bo Zeng, Hao Wang, Tianqi Shi, Liangying Shao, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. 2025. https://arxiv.org/abs/2503.01461 Towards widening the distillation bottleneck for reasoning models . Preprint, arXiv:2503.01461
2025 arXiv
-
[98]
Zhaojian Yu, Yinghao Wu, Yilun Zhao, Arman Cohan, and Xiao-Ping Zhang. 2025. https://arxiv.org/abs/2504.00810 Z1: Efficient test-time scaling with code . Preprint, arXiv:2504.00810
2025 arXiv
-
[99]
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, and Gao Huang. 2025. https://arxiv.org/abs/2504.13837 Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model? Preprint, arXiv:2504.13837
2025 arXiv
-
[100]
Michał Zawalski, William Chen, Karl Pertsch, Oier Mees, Chelsea Finn, and Sergey Levine. 2025. https://arxiv.org/abs/2407.08693 Robotic control via embodied chain-of-thought reasoning . Preprint, arXiv:2407.08693
2025 arXiv
-
[101]
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476--15488
2022
-
[102]
Qiyuan Zhang, Fuyuan Lyu, Zexu Sun, Lei Wang, Weixu Zhang, Wenyue Hua, Haolun Wu, Zhihan Guo, Yufei Wang, Niklas Muennighoff, Irwin King, Xue Liu, and Chen Ma. 2025. https://arxiv.org/abs/2503.24235 A survey on test-time scaling in large language models: What, how, where, and ...
2025 arXiv
-
[103]
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. https://arxiv.org/abs/2210.03493 Automatic chain of thought prompting in large language models . Preprint, arXiv:2210.03493
2022 arXiv
-
[104]
Wanjun Zhong, Siyuan Wang, Duyu Tang, Zenan Xu, Daya Guo, Jiahai Wang, Jian Yin, Ming Zhou, and Nan Duan. 2021. https://arxiv.org/abs/2104.06598 Ar-lsat: Investigating analytical reasoning of text . Preprint, arXiv:2104.06598
2021 arXiv
-
[105]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[106]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.