Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that prefix-tuning a 1.5B math reasoning model for diversity reaches 80% accuracy with 32 samples rather than 256.

desk verdict A competent survey plus a small, honest empirical study whose central diversity claim is not yet supported by the evidence, and which contains a clear internal inconsistency in the reported numbers. read the letter →

arxiv 2506.04611 v1 pith:KOO3ODKE submitted 2025-06-05 cs.CL

classification cs.CL
keywords test-timescalinggenerativediversityprefixtuningbest-of-Nsamplingmathematicalreasoningdistillationmajorityvoting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Test-time scaling works only if the samples a model draws at inference actually differ; this paper argues that reasoning-optimized, distilled models lose that variety and therefore scale poorly with best-of-N sampling. To fix that, the authors propose ADAPT, a lightweight prefix fine-tuning method that mixes diverse Qwen2.5-Math-1.5B responses into the target model's own outputs and tunes only the first 512 tokens of the reasoning trajectory. On mathematical reasoning tasks, ADAPT reaches 80% majority-vote accuracy with 32 samples, eight times fewer than the 256 samples the unmodified distilled baseline needs. If the diversity hypothesis is right, diversity-aware tuning is a cheap way to make test-time scaling practical for small reasoning models.

What carries the argument

The load-bearing mechanism is prefix fine-tuning over a curated data mixture. The training set is 90% responses from Qwen2.5-Math-1.5B generated with a prompt that asks for the initial step toward solving the question, which is meant to encourage varied reasoning beginnings, plus 10% of the target model's own outputs to prevent forgetting. All examples are truncated to 512 tokens and only prefix parameters are updated while the rest of the model stays frozen; at inference, the tuned model is evaluated with best-of-N sampling and majority voting.

What would settle it

Compute self-BLEU and pairwise entropy over the N=32 candidate solutions from ADAPT and from DeepSeek-R1-Distill-Qwen-1.5B at temperature 0.8; if ADAPT's accuracy advantage appears while its diversity metrics are no higher, the paper's mechanism is wrong. A control trained on 90% Qwen responses without the initial-step prompt that still gains accuracy would likewise show that transfer, not diversity, drives the result.

Watch

Extended reading notes

Core claim

On its own terms, the paper shows that a small prefix fine-tune can restore the output variety that reasoning distillation removes, and that this variety is what makes sampling-based test-time scaling pay off. With only the first 512 tokens of each training example and a 90/10 mixture of Qwen2.5-Math-1.5B and DeepSeek-R1-Distill-Qwen-1.5B outputs, ADAPT reaches 80% accuracy at N=32, versus N=256 for the unmodified distilled model, and peaks at 81.0%. The paper takes this as evidence that generative diversity is a key enabler of test-time scaling, not just a side effect of sampling temperature.

Load-bearing premise

The paper assumes its 90/10 training mixture increases the variety of solution paths the target model produces, and that the accuracy gains come from that variety rather than from borrowed reasoning ability, yet it never measures diversity directly.

Editorial extensions

If this is right

  • Reasoning distillation, while raising single-sample accuracy, can quietly suppress the output diversity that sampling-based test-time scaling depends on.
  • ADAPT reaches 80% accuracy at N=32 rather than N=256, an 8x reduction in the inference budget needed to hit the same threshold.
  • Most of ADAPT's gains arrive before N=32, so small compute budgets capture most of the benefit of test-time scaling.
  • At N=256, ADAPT's peak accuracy of 81.0% slightly exceeds the distilled baseline's 80.8%, so the efficiency gain is not bought with peak performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If diversity is the active ingredient, the same prefix-tuning recipe should transfer to other low-diversity reasoning models and other aggregation rules such as verifier reranking; the paper only tests majority voting on math.
  • A cleaner test would compare ADAPT against a control trained on 90% Qwen responses with the standard chat template, isolating the contribution of the initial-step prompt from raw knowledge transfer.
  • Because only prefix parameters are updated, the method should be cheap enough to apply at larger scale, but the paper does not test this directly.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper combines a survey of test-time scaling (TTS) methods—taxonomized into sampling, search, and trajectory optimization—with a proposed method, ADAPT, that applies prefix tuning to a 1.5B distilled reasoning model using a 90/10 mixture of Qwen2.5-Math-1.5B and DeepSeek-R1-Distill-Qwen-1.5B outputs. The empirical claim is that ADAPT reaches 80% accuracy with N=32 best-of-N majority-vote samples, an 8x efficiency gain over the DeepSeek baseline that requires N=256, and that this gain stems from increased generative diversity. The paper also includes future-directions discussion and an appendix with training details.

Significance. If the central claim holds, the paper would provide a simple and parameter-efficient way to improve test-time scaling in small reasoning-optimized models, with potentially broad applicability. The survey component offers a useful, if not deeply novel, organization of recent TTS literature. The authors are explicitly honest about limitations, and the reported efficiency numbers, if reproducible, are practically interesting. However, the key attribution of the gain to diversity is not directly supported by any measurement, and the experimental evidence rests on a single run on an unnamed benchmark. Consequently, the significance is conditional on additional validation that is within the scope of a revision.

major comments (4)
  1. [Section 5.4, Table 1] The text states that 'At N=16, it already surpasses the distilled baseline at N=256,' but Table 1 reports ADAPT at 78.2% and DeepSeek Qwen-1.5B at 80.8% at those respective sample counts. The claim is contradicted by the paper's own table and must be corrected or substantiated with the actual accuracy at N=16.
  2. [Section 7 (Limitations), Section 4.1] The paper acknowledges that it does not directly measure diversity (e.g., self-BLEU, pairwise entropy) and that the diversity effect is 'inferred only through indirect accuracy gains.' This is load-bearing because the title, method name, and Section 5.5 all attribute ADAPT's efficiency improvement to increased diversity, yet the training mixture is 90% Qwen2.5 responses, making content transfer or answer-format transfer a live confound. Without any diversity metric computed on generated outputs, the accuracy curves in Table 1 cannot distinguish the diversity hypothesis from a content-transfer hypothesis. The revision should measure diversity directly (e.g., self-BLEU, n-gram overlap, or pairwise entropy of sampled solutions at a given N) or explicitly soften the causal claim.
  3. [Section 5.1, Experiments] The evaluation benchmark is described only as 'akin to MATH-500' without naming the exact dataset, and there is no mention of random seeds, multiple runs, or variance estimates. For a paper whose central claim is a quantitative efficiency ratio (8x) derived from a single accuracy threshold, this is insufficient experimental reporting. The authors should specify the exact benchmark, provide results over multiple seeds with standard errors or confidence intervals, and report the actual accuracy at N=32 (the text infers Min N but does not show the measured value).
  4. [Section 4.1, Section 5.4] The 90/10 mixture ratio and the custom prompt template are free parameters of the method, yet no ablations are performed to isolate the contribution of each component. Without ablations (e.g., 100% DeepSeek data with the same prefix tuning, or 100% Qwen data without the prompt template), it is not possible to tell whether the observed gains are due to diversity-oriented data selection, the prompt template, or simply the added Qwen-derived training signal. Reporting such ablations would also help test the diversity mechanism.
minor comments (5)
  1. [Abstract, Section 5.4] The abstract says 'eight times less compute' while the body says '8× improvement' in sample count; please use consistent wording and clarify that the comparison is in number of samples, not wall-clock compute.
  2. [Figure 5 caption] The caption reads 'Marginal gain per generation for across models' and should be corrected to '...for all models'.
  3. [Section A.5.2] The description of the dataset split ('combined dataset is shuffled and split into 90% for training and 10% for testing') is ambiguous relative to Section 4.1's 90/10 mixture of model outputs; clarify whether the 90/10 refers to data source mixture or train/test split.
  4. [Appendix A.1, Figure 6] Several entries in the methodology structure are duplicated (e.g., Coconut, Latent-Thought, CODI, and Looped transformer appear under both 'Search/Hidden layer search' and 'Search/Self-improvement'); these should be listed under the most appropriate category only.
  5. [References] Some references are incomplete or nonstandard: 'Beeching et al.' appears without a year or full citation, 'Chen et al.' in the Safety paragraph of Section 6 lacks a reference, and several entries contain 'and 1 others' instead of the full author list. Please review and complete the bibliography.

Circularity Check

1 steps flagged · score 2.0 of 10

No circularity in the efficiency measurement; the diversity attribution is a partly self-referential interpretation, acknowledged in the Limitations, so only a low score.

  1. other [Section 5.5 and Limitations (p. 8-9); also Sections 4.1 and 5.3]
    "Although ADAPT improves sample efficiency, it does not directly optimize diversity metrics (e.g., self-BLEU, pairwise entropy), and its diversity-enhancing effect is inferred only through indirect accuracy gains. ... Our results validate the central hypothesis: enhancing output diversity improves both the accuracy and efficiency of reasoning models under Best-of-N sampling."

    The paper's central interpretive claim is that ADAPT's gains validate the diversity hypothesis, but ADAPT's diversity enhancement is never measured. The training mixture is labeled 'diverse' because Qwen2.5-Math-1.5B is assumed to have higher generative diversity, and that assumption is supported in Section 5.3 by the same TTS scaling behavior the hypothesis is meant to explain. The accuracy gains are then used to confirm the diversity hypothesis, while the diversity increase is inferred only from those accuracy gains. This closes an interpretive loop: diversity is inferred from the accuracy improvement, and the accuracy improvement is explained by diversity. The headline empirical efficiency comparison is not circular, which is why the score is low.

full rationale

The headline result—ADAPT reaches 80% accuracy at N=32 versus N=256 for DeepSeek-Qwen-1.5B—is an empirical measurement against external baselines under a fixed Best-of-N majority-voting protocol. It is not derived by fitting or by construction, so the central efficiency claim is self-contained. The only circular element is interpretive: the claim that the results validate the diversity hypothesis rests on the unmeasured assumption that the 90% Qwen training mixture increases diversity. The Limitations section explicitly concedes that no direct diversity metric is measured and that the diversity effect is 'inferred only through indirect accuracy gains.' This makes the diversity explanation self-referential, but it does not invalidate the measured accuracy comparison. Score 2 reflects a minor interpretive circularity, not a construction-level circularity.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new entities. Its central claim depends on several hand-chosen hyperparameters and on the assumption that the training mixture increases diversity without direct measurement. The most honest measure is that the method is a standard prefix-tuning recipe applied to a novel data mixture.

free parameters (4)
  • data_mixture_ratio = 90% Qwen / 10% DeepSeek
    Chosen by hand without a sweep; the paper states this mixture in Section 4.1.
  • prefix_truncation_length = 512 tokens
    Fixed to 512 tokens for all training instances; no analysis of the effect of this cutoff.
  • training_epochs = 3
    Reported in Appendix A.5.2; no tuning or justification.
  • learning_rate = 5e-6
    Reported in Appendix A.5.2; no sweep.
assumptions (3)
  • domain assumption Majority voting over N samples is a valid test-time scaling strategy and accuracy improves with N.
    Used as the evaluation protocol in Section 5.1; the paper does not justify why this is the right TTS setup.
  • domain assumption Qwen2.5-Math-1.5B responses with the custom 'initial step' prompt are a valid source of diverse reasoning prefixes.
    This is the foundation of the dataset in Section 4.1; no validation that the prompt actually increases diversity.
  • domain assumption The evaluation benchmark is representative of mathematical reasoning despite not being named.
    Section 5.1 says 'akin to MATH-500' but does not identify the actual benchmark, so the reader must assume the result transfers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning." pith.science (2026). https://pith.science/paper/KOO3ODKE

@misc{pith2026250604611,
  author       = {Pith},
  title        = {Pith review of: Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KOO3ODKE}},
  note         = {Machine review of arXiv:2506.04611}
}
read the original abstract

Test-Time Scaling (TTS) improves the reasoning performance of Large Language Models (LLMs) by allocating additional compute during inference. We conduct a structured survey of TTS methods and categorize them into sampling-based, search-based, and trajectory optimization strategies. We observe that reasoning-optimized models often produce less diverse outputs, which limits TTS effectiveness. To address this, we propose ADAPT (A Diversity Aware Prefix fine-Tuning), a lightweight method that applies prefix tuning with a diversity-focused data strategy. Experiments on mathematical reasoning tasks show that ADAPT reaches 80% accuracy using eight times less compute than strong baselines. Our findings highlight the essential role of generative diversity in maximizing TTS effectiveness.

Figures

Figures reproduced from arXiv: 2506.04611 by the authors.

Figure 1
Figure 1. Accuracy vs. Inference Cost (log scale). Each point represents a language model. While TTS has shown effectiveness, its perfor￾mance is often tied to the model’s intrinsic capac￾ity for generation diversity—a factor not yet well understood or explicitly optimized. In particular, models optimized for reasoning, such as distilled variants, tend to exhibit reduced output variance, which may dampen the gains from TTS. T… view at source ↗
Figure 2
Figure 2. Sampling-based method. The model sam￾ples N candidates, a verifier scores each, and the highest-scoring answer is returned. Some variants re￾peat this loop for multiple rounds. and proposed various strategies to improve perfor￾mance. For example, Tian et al. (2025) proposed a method that generates answers in multiple rounds. In each round, the model uses the previous answer and the original input as a new input. Thi… view at source ↗
Figure 3
Figure 3. Search-Based The schematic illustrates various search-based methods, including Chain of Thought (CoT), Self-Consistency with Chain of Thought (CoT-SC), Tree of Thoughts (ToT), Graph of Thoughts (GoT), Forest of Thoughts (FoT), and Atom of Thoughts (AoT). generate self-criticism and revise their responses accordingly. Building on this, STaR (Zelikman et al., 2022) bootstraps high-quality rationales that lead to corre… view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: Marginal gain per generation for across models only N = 16 samples, ADAPT already achieves 78.2%, making it suitable for low-budget inference. In contrast, Qwen2.5-1.5B fails to reach the 80% threshold even at N = 256 [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Methodology classification structure. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling

    cs.CV 2026-08 conditional novelty 7.0 of 10

    Perturbation-based selection does not beat a format-matched control that spends the same short-answer budget on the unperturbed image, so reported gains against CoT-only majority voting are mostly a decoding-format effect.

Reference graph

Works this paper leans on

106 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. https://arxiv.org/abs/2108.07732 Program synthesis with large language models . Preprint, arXiv:2108.07732

  2. [2]

    https://huggingface.co/spaces/HuggingFaceH4/blogpost-scaling-test-time-compute Scaling test-time compute with open models

    Edward Beeching, Lewis Tunstall, and Sasha Rush. https://huggingface.co/spaces/HuggingFaceH4/blogpost-scaling-test-time-compute Scaling test-time compute with open models

  3. [3]

    Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2024. https://doi.org/10.1609/aaai.v38i16.29720 Graph of thoughts: Solving elaborate problems with large language models . Proceedings of the AAAI Conference on Artificial ...

  4. [4]

    Zhenni Bi, Kai Han, Chuanjian Liu, Yehui Tang, and Yunhe Wang. 2025. https://arxiv.org/abs/2412.09078 Forest-of-thought: Scaling test-time compute for enhancing llm reasoning . Preprint, arXiv:2412.09078

  5. [5]

    Houda Bouamor, Juan Pino, and Kalika Bali, editors. 2023. https://aclanthology.org/2023.emnlp-main.0/ Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Singapore

  6. [6]

    Feng Chen, Allan Raventos, Nan Cheng, Surya Ganguli, and Shaul Druckmann. 2025 a . https://arxiv.org/abs/2502.07154 Rethinking fine-tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning . Preprint, arXiv:2502.07154

  7. [7]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, and 39 others. 2021. https://arxiv.org/abs/2107.03374 Evaluating large lang...

  8. [8]

    Xinghao Chen, Zhijing Sun, Wenjin Guo, Miaoran Zhang, Yanjun Chen, Yirong Sun, Hui Su, Yijie Pan, Dietrich Klakow, Wenjie Li, and Xiaoyu Shen. 2025 b . https://arxiv.org/abs/2502.18001 Unveiling the key factors for distilling chain-of-thought reasoning . Preprint, arXiv:2502.18001

Show all 106 references
  1. [9]

    Reasoning models don’t always say what they think

    Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner Fabien Roger Vlad Mikulik, Sam Bowman, Jan Leike Jared Kaplan, and 1 others. Reasoning models don’t always say what they think

  2. [10]

    Jie Cheng, Ruixi Qiao, Lijun Li, Chao Guo, Junle Wang, Gang Xiong, Yisheng Lv, and Fei-Yue Wang. 2025. https://arxiv.org/abs/2504.15275 Stop summation: Min-form credit assignment is all process reward model needs for reasoning . Preprint, arXiv:2504.15275

  3. [11]

    Yinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang, Bo Dai, Sridhar Thiagarajan, Craig Boutilier, Rishabh Agarwal, Aviral Kumar, and Aleksandra Faust. 2024. https://arxiv.org/abs/2412.15287 Inference-aware fine-tuning for best-of-n sampling in large language models . P...

  4. [12]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, and 1 others. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311

  5. [13]

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://arxiv.org/abs/1803.05457 Think you have solved question answering? try arc, the ai2 reasoning challenge . Preprint, arXiv:1803.05457

  6. [14]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . ...

  7. [15]

    Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Chun Cheung, Jan Kautz, Carlos Guestrin, Tatsunori Hashimoto, Sanmi Koyejo, Yejin Choi, Yu Sun, and Xiaolong Wang. 2025. https://arxiv.org/abs/2504.05298 One-minute video generation wi...

  8. [16]

    Xingyu Dang, Christina Baek, Kaiyue Wen, Zico Kolter, and Aditi Raghunathan. 2025. https://arxiv.org/abs/2504.10478 Weight ensembling improves reasoning in language models . Preprint, arXiv:2504.10478

  9. [17]

    DeepSeek-AI. 2025. https://arxiv.org/abs/2501.12948 Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning . Preprint, arXiv:2501.12948

  10. [18]

    Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. 2023. https://arxiv.org/abs/2305.15408 Towards revealing the mystery behind chain of thought: A theoretical perspective . Preprint, arXiv:2305.15408

  11. [19]

    Sicheng Feng, Gongfan Fang, Xinyin Ma, and Xinchao Wang. 2025. https://arxiv.org/abs/2504.10903 Efficient reasoning models: A survey . Preprint, arXiv:2504.10903

  12. [20]

    Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang. 2024. https://arxiv.org...

  13. [21]

    Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. https://arxiv.org/abs/2211.10435 Pal: Program-aided language models . Preprint, arXiv:2211.10435

  14. [22]

    Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. https://doi.org/10.1162/tacl_a_00370 Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies . Transactions of the Association for Computational Lingu...

  15. [23]

    Anna Goldie, Azalia Mirhoseini, Hao Zhou, Irene Cai, and Christopher D. Manning. 2025. https://arxiv.org/abs/2504.04736 Synthetic data generation & multi-step rl for reasoning & tool use . Preprint, arXiv:2504.04736

  16. [24]

    Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. 2024. https://arxiv.org/abs/2412.06769 Training large language models to reason in a continuous latent space . Preprint, arXiv:2412.06769

  17. [25]

    Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. 2024. https://arxiv.org/abs/2402.14008 Olympiadbench: A challenging benchmark for promoting agi with olym...

  18. [26]

    Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021 a . Measuring coding challenge competence with apps. NeurIPS

  19. [27]

    Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021 b . Aligning ai with shared human values. Proceedings of the International Conference on Learning Representations (ICLR)

  20. [28]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 c . Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR)

  21. [29]

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 d . Measuring mathematical problem solving with the math dataset. NeurIPS

  22. [30]

    Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, and Matthias Bethge. 2025. https://arxiv.org/abs/2504.07086 A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility . Preprint, arXiv:2504.07086

  23. [31]

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osin...

  24. [32]

    Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023. https://arxiv.org/abs/2305.02301 Distilling step-by-step! outperforming larger language models with less training data and smal...

  25. [33]

    Audrey Huang, Adam Block, Qinghua Liu, Nan Jiang, Akshay Krishnamurthy, and Dylan J. Foster. 2025 a . https://arxiv.org/abs/2503.21878 Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment . Preprint, arXiv:2503.21878

  26. [34]

    Chengsong Huang, Langlin Huang, Jixuan Leng, Jiacheng Liu, and Jiaxin Huang. 2025 b . https://arxiv.org/abs/2503.00031 Efficient test-time scaling via self-calibration . Preprint, arXiv:2503.00031

  27. [35]

    Jianhao Huang, Zixuan Wang, and Jason D. Lee. 2025 c . https://arxiv.org/abs/2502.21212 Transformers learn to implement multi-step gradient descent with chain of thought . Preprint, arXiv:2502.21212

  28. [36]

    Xiaoke Huang, Juncheng Wu, Hui Liu, Xianfeng Tang, and Yuyin Zhou. 2025 d . https://arxiv.org/abs/2504.00869 m1: Unleash the potential of test-time scaling for medical reasoning with large language models . Preprint, arXiv:2504.00869

  29. [37]

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024. https://arxiv.org/abs/2403.07974 Livecodebench: Holistic and contamination free evaluation of large language models for code . Preprint, a...

  30. [38]

    Metaxas, and Tong Che

    Can Jin, Hongwu Peng, Qixin Zhang, Yujin Tang, Dimitris N. Metaxas, and Tong Che. 2025. https://arxiv.org/abs/2504.09772 Two heads are better than one: Test-time scaling of multi-agent collaborative reasoning . Preprint, arXiv:2504.09772

  31. [39]

    Manuj Kant, Manav Kant, Marzieh Nabi, Preston Carlson, and Megan Ma. 2024. https://arxiv.org/abs/2410.09904 Equitable access to justice: Logical llms show promise . Preprint, arXiv:2410.09904

  32. [40]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. https://arxiv.org/abs/2001.08361 Scaling laws for neural language models . Preprint, arXiv:2001.08361

  33. [41]

    Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. https://doi.org/10.18653/v1/N16-1136 MAWPS : A math word problem repository . In Proceedings of the 2016 Conference of the North A merican Chapter of the Association for Computational L...

  34. [42]

    Deqian Kong, Minglu Zhao, Dehong Xu, Bo Pang, Shu Wang, Edouardo Honig, Zhangzhang Si, Chuan Li, Jianwen Xie, Sirui Xie, and Ying Nian Wu. 2025. https://arxiv.org/abs/2502.01567 Scalable language models with posterior inference of latent thought vectors . Preprint, arXiv:2502.01567

  35. [43]

    Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal Behbahani, and Aleksandra Faust. 2025...

  36. [44]

    Komal Kumar, Tajamul Ashraf, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, Phillip H. S. Torr, Fahad Shahbaz Khan, and Salman Khan. 2025 b . https://arxiv.org/abs/2502.21321 Llm post-training: A deep dive into reasoning large language mod...

  37. [45]

    Fangyu Lei, Qian Liu, Yiming Huang, Shizhu He, Jun Zhao, and Kang Liu. 2024. https://arxiv.org/abs/2310.15147 S3eval: A synthetic, scalable, systematic evaluation suite for large language models . Preprint, arXiv:2310.15147

  38. [46]

    Xinzhe Li. 2025. https://openreview.net/forum?id=x9VQFjtOPS A survey on LLM test-time compute via search: Tasks, LLM profiling, search algorithms, and relevant frameworks . Transactions on Machine Learning Research

  39. [47]

    Yanyang Li, Michael Lyu, and Liwei Wang. 2025. Learning to reason from feedback at test-time

  40. [48]

    Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma. 2024. https://arxiv.org/abs/2402.12875 Chain of thought empowers transformers to solve inherently serial problems . Preprint, arXiv:2402.12875

  41. [49]

    Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let's verify step by step. arXiv preprint arXiv:2305.20050

  42. [50]

    Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. https://openreview.net/forum?id=v8L0pN6EOi Let's verify step by step . In The Twelfth International Conference on Learning Re...

  43. [51]

    Gonzalez

    Kevin Lin, Charlie Snell, Yu Wang, Charles Packer, Sarah Wooders, Ion Stoica, and Joseph E. Gonzalez. 2025. https://arxiv.org/abs/2504.13171 Sleep-time compute: Beyond inference scaling at test-time . Preprint, arXiv:2504.13171

  44. [52]

    Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. https://doi.org/10.18653/v1/P17-1015 Program induction by rationale generation: Learning to solve and explain algebraic word problems . In Proceedings of the 55th Annual Meeting of the Association for Computational ...

  45. [53]

    Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang

    Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023. https://arxiv.org/abs/2210.10749 Transformers learn shortcuts to automata . Preprint, arXiv:2210.10749

  46. [54]

    Guanlin Liu, Anand Ramachandran, Tanmay Gangwani, Yan Fu, and Abhinav Sethy. 2025. https://arxiv.org/abs/2502.17717 Knowledge distillation with training wheels . Preprint, arXiv:2502.17717

  47. [55]

    Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

    Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica. 2025. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl. https://pretty-radio-b75.notion.site/DeepScaleR-Su...

  48. [56]

    Junyu Ma, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, and Dong Yu. 2025 a . https://arxiv.org/abs/2505.03320 Recall with reasoning: Chain-of-thought distillation for mamba's long-context memory and extrapolation . Preprint, arXiv:2505.03320

  49. [57]

    Xiao Ma, Yuhui Tao, Yuhan Zhang, Zexuan Ji, Yizhe Zhang, and Qiang Chen. 2024. https://arxiv.org/abs/2406.17608 Test-time generative augmentation for medical image segmentation . Preprint, arXiv:2406.17608

  50. [58]

    Yingwei Ma, Yongbin Li, Yihong Dong, Xue Jiang, Rongyu Cao, Jue Chen, Fei Huang, and Binhua Li. 2025 b . https://arxiv.org/abs/2503.23803 Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute . Preprint, arXiv:2503.23803

  51. [59]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, and 1 others. 2023. Self-refine: Iterative refinement with self-feedback

  52. [60]

    William Merrill and Ashish Sabharwal. 2024. https://arxiv.org/abs/2310.07923 The expressive power of transformers with chain of thought . Preprint, arXiv:2310.07923

  53. [61]

    Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. 2024. https://arxiv.org/abs/2410.05229 Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models . Preprint, arXiv:2410.05229

  54. [62]

    NVIDIA, :, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, and 60 others. 202...

  55. [63]

    OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  56. [64]

    OpenAI . 2024 a . https://openai.com/index/learning-to-reason-with-llms/ Learning to reason with llms . Accessed: 2025-05-18

  57. [65]

    OpenAI . 2024 b . https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf Openai o3 and o4-mini system card . Technical report, OpenAI. Accessed: 2025-04-22

  58. [66]

    Qianjun Pan, Wenkai Ji, Yuyang Ding, Junsong Li, Shilian Chen, Junyi Wang, Jie Zhou, Qin Chen, Min Zhang, Yulan Wu, and Liang He. 2025. https://arxiv.org/abs/2505.02665 A survey of slow thinking-based reasoning llms using reinforced learning and inference-time scaling law . Pr...

  59. [67]

    Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. https://doi.org/10.18653/v1/2021.naacl-main.168 Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Ling...

  60. [68]

    Yuxiao Qu, Matthew Y. R. Yang, Amrith Setlur, Lewis Tunstall, Edward Emanuel Beeching, Ruslan Salakhutdinov, and Aviral Kumar. 2025. https://arxiv.org/abs/2503.07572 Optimizing test-time compute via meta reinforcement fine-tuning . Preprint, arXiv:2503.07572

  61. [69]

    Yuxiao Qu, Tianjun Zhang, Naman Garg, and Aviral Kumar. 2024. Recursive introspection: Teaching language model agents how to self-improve

  62. [70]

    Leonardo Ranaldi, Marco Valentino, Alexander Polonsky, and Andrè Freitas. 2025. https://arxiv.org/abs/2502.12616 Improving chain-of-thought reasoning via quasi-symbolic abstractions . Preprint, arXiv:2502.12616

  63. [71]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2024. https://openreview.net/forum?id=Ti67584b98 GPQA : A graduate-level google-proof q&a benchmark . In First Conference on Language Modeling

  64. [72]

    Abulhair Saparov and He He. 2023. https://openreview.net/forum?id=qFVVBzXxR2V Language models are greedy reasoners: A systematic formal analysis of chain-of-thought . In The Eleventh International Conference on Learning Representations

  65. [73]

    Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J. Reddi. 2025. https://arxiv.org/abs/2502.17416 Reasoning with latent thoughts: On the power of looped transformers . Preprint, arXiv:2502.17416

  66. [74]

    Amrith Setlur, Nived Rajaraman, Sergey Levine, and Aviral Kumar. 2025. https://arxiv.org/abs/2502.12118 Scaling test-time compute without verification or rl is suboptimal . Preprint, arXiv:2502.12118

  67. [75]

    Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, and Yulan He. 2025. https://arxiv.org/abs/2502.21074 Codi: Compressing chain-of-thought into continuous space via self-distillation . Preprint, arXiv:2502.21074

  68. [76]

    Kumar Shridhar, Alessandro Stolfo, and Mrinmaya Sachan. 2023. https://arxiv.org/abs/2212.00193 Distilling reasoning capabilities into smaller language models . Preprint, arXiv:2212.00193

  69. [77]

    Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. https://arxiv.org/abs/2408.03314 Scaling llm test-time compute optimally can be more effective than scaling model parameters . Preprint, arXiv:2408.03314

  70. [78]

    Benedikt Stroebl, Sayash Kapoor, and Arvind Narayanan. 2024. https://arxiv.org/abs/2411.17501 Inference scaling flaws: The limits of llm resampling with imperfect verifiers . Preprint, arXiv:2411.17501

  71. [79]

    Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and Xia Hu. 2025. https://arxiv.org/abs/2503.16419 Stop overthinking: A survey on efficient reasoning for large language models . Preprint, arXiv...

  72. [80]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/v1/N19-1421 C ommonsense QA : A question answering challenge targeting commonsense knowledge . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associa...

  73. [81]

    Fengwei Teng, Zhaoyang Yu, Quan Shi, Jiayi Zhang, Chenglin Wu, and Yuyu Luo. 2025. https://arxiv.org/abs/2502.12018 Atom of thoughts for markov llm test-time scaling . Preprint, arXiv:2502.12018

  74. [82]

    Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yunjie Ji, Yiping Peng, Han Zhao, and Xiangang Li. 2025. https://arxiv.org/abs/2503.19855 Think twice: Enhancing llm reasoning by scaling multi-round test-time thinking . Preprint, arXiv:2503.19855

  75. [83]

    Guiyao Tie, Zeli Zhao, Dingjie Song, Fuyang Wei, Rong Zhou, Yurou Dai, Wen Yin, Zhejian Yang, Jiangyue Yan, Yao Su, Zhenhan Dai, Yifeng Xie, Yihan Cao, Lichao Sun, Pan Zhou, Lifang He, Hechang Chen, Yu Zhang, Qingsong Wen, and 7 others. 2025. https://arxiv.org/abs/2503.06072 A...

  76. [84]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, and 1 others. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971

  77. [85]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30

  78. [86]

    Jiankang Wang, Jianjun Xu, Xiaorui Wang, Yuxin Wang, Mengting Xing, Shancheng Fang, Zhineng Chen, Hongtao Xie, and Yongdong Zhang. 2025 a . https://arxiv.org/abs/2412.08864 A graph-based synthetic data pipeline for scaling high-quality reasoning instructions . Preprint, arXiv:...

  79. [87]

    Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou. 2024 a . https://arxiv.org/abs/2406.04692 Mixture-of-agents enhances large language model capabilities . Preprint, arXiv:2406.04692

  80. [88]

    Rui Wang, Hongru Wang, Boyang Xue, Jianhui Pang, Shudong Liu, Yi Chen, Jiahao Qiu, Derek Fai Wong, Heng Ji, and Kam-Fai Wong. 2025 b . https://arxiv.org/abs/2503.24377 Harnessing the reasoning economy: A survey of efficient reasoning for large language models . Preprint, arXiv...

  81. [89]

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. https://arxiv.org/abs/2203.11171 Self-consistency improves chain of thought reasoning in language models . Preprint, arXiv:2203.11171

  82. [90]

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. 2024 b . https://arxiv.org/abs/2406.01574 Mmlu-pro: A more robust ...

  83. [91]

    Sean Welleck, Amanda Bertsch, Matthew Finlayson, Hailey Schoelkopf, Alex Xie, Graham Neubig, Ilia Kulikov, and Zaid Harchaoui. 2024. https://arxiv.org/abs/2406.16838 From decoding to meta-generation: Inference-time algorithms for large language models . Preprint, arXiv:2406.16838

  84. [92]

    Lillicrap, Kenji Kawaguchi, and Michael Shieh

    Yuxi Xie, Anirudh Goyal, Wenyue Zheng, Min-Yen Kan, Timothy P. Lillicrap, Kenji Kawaguchi, and Michael Shieh. 2024. https://arxiv.org/abs/2405.00451 Monte carlo tree search boosts reasoning via iterative preference learning . Preprint, arXiv:2405.00451

  85. [93]

    Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, Chenyang Shao, Yuwei Yan, Qinglong Yang, Yiwen Song, Sijian Ren, Xinyuan Hu, Yu Li, Jie Feng, Chen Gao, and Yong Li. 2025 a . https://arxiv.or...

  86. [94]

    Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He. 2025 b . https://arxiv.org/abs/2502.18600 Chain of draft: Thinking faster by writing less . Preprint, arXiv:2502.18600

  87. [95]

    Wenkai Yang, Shuming Ma, Yankai Lin, and Furu Wei. 2025. https://arxiv.org/abs/2502.18080 Towards thinking-optimal scaling of test-time compute for llm reasoning . Preprint, arXiv:2502.18080

  88. [96]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. https://arxiv.org/abs/2305.10601 Tree of thoughts: Deliberate problem solving with large language models . Preprint, arXiv:2305.10601

  89. [97]

    Huifeng Yin, Yu Zhao, Minghao Wu, Xuanfan Ni, Bo Zeng, Hao Wang, Tianqi Shi, Liangying Shao, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. 2025. https://arxiv.org/abs/2503.01461 Towards widening the distillation bottleneck for reasoning models . Preprint, arXiv:2503.01461

  90. [98]

    Zhaojian Yu, Yinghao Wu, Yilun Zhao, Arman Cohan, and Xiao-Ping Zhang. 2025. https://arxiv.org/abs/2504.00810 Z1: Efficient test-time scaling with code . Preprint, arXiv:2504.00810

  91. [99]

    Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, and Gao Huang. 2025. https://arxiv.org/abs/2504.13837 Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model? Preprint, arXiv:2504.13837

  92. [100]

    Michał Zawalski, William Chen, Karl Pertsch, Oier Mees, Chelsea Finn, and Sergey Levine. 2025. https://arxiv.org/abs/2407.08693 Robotic control via embodied chain-of-thought reasoning . Preprint, arXiv:2407.08693

  93. [101]

    Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476--15488

  94. [102]

    Qiyuan Zhang, Fuyuan Lyu, Zexu Sun, Lei Wang, Weixu Zhang, Wenyue Hua, Haolun Wu, Zhihan Guo, Yufei Wang, Niklas Muennighoff, Irwin King, Xue Liu, and Chen Ma. 2025. https://arxiv.org/abs/2503.24235 A survey on test-time scaling in large language models: What, how, where, and ...

  95. [103]

    Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. https://arxiv.org/abs/2210.03493 Automatic chain of thought prompting in large language models . Preprint, arXiv:2210.03493

  96. [104]

    Wanjun Zhong, Siyuan Wang, Duyu Tang, Zenan Xu, Daya Guo, Jiahai Wang, Jian Yin, Ming Zhou, and Nan Duan. 2021. https://arxiv.org/abs/2104.06598 Ar-lsat: Investigating analytical reasoning of text . Preprint, arXiv:2104.06598

  97. [105]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  98. [106]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.