Pith. sign in

Paper Citation Record · LEDGER

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

As of 17 August 2026, this Paper Citation Record lists 100 of 283 outbound references and 0 inbound Pith citation observations for arXiv:2608.04001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04001 v1

Coverage vector

measured 100 of 283 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:49:04.602845Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 283 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved95
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 716c193e-8f4e-4fe9-b004-ee536ac34a12 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Curious Case of Neural Text Degeneration

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.053337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.053337Z digest=sha256:bf209a3674e1ffb11f683965be29cf40a5b373ad63b91198fd0870dd00f23b43

Observation 574b94b4-bd18-42be-989f-9f0fdd92393f · outbound

This paper cites Optimizing Large Language Model Hyperparameters for Code Generation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Optimizing Large Language Model Hyperparameters for Code Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.064748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.064748Z digest=sha256:439416cfeca31d944d815ad8a85b7417404700c7133c189a21d825c0fed87a23

Observation 2ec0beb8-9236-414e-bf08-9e8fec269b03 · outbound

This paper cites arXiv preprint arXiv:2407.01082 , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility arXiv preprint arXiv:2407.01082 , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.078009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.078009Z digest=sha256:5bd17e01825783acc11c8548e9f9861fd00e2778cb64d000ffa58d0d01281ef4

Observation 5211dacc-8084-4023-b3be-f4ce55c42205 · outbound

This paper cites Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.090744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.090744Z digest=sha256:aa4c8956f2629ba57e9e2df05784be60a2d1cfae2ca38db1aa96dc4da52ab139

Observation cfe11679-bd56-40d2-8b17-5901f17e2972 · outbound

This paper cites A Thorough Examination of Decoding Methods in the Era of LLMs.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility A Thorough Examination of Decoding Methods in the Era of LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.109318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.109318Z digest=sha256:daeee03aff0ff229f5ca44d914f4ecee2058ddb578d9fb603c824118addf4c2a

Observation c69370a4-72d4-494f-9e74-37073bd86801 · outbound

This paper cites Closing the Curious Case of Neural Text Degeneration.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Closing the Curious Case of Neural Text Degeneration

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.120745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.120745Z digest=sha256:dcf3679b08b5d4d488c2fd8d06d49185fd2fbffa50ff3396d3bf6ece4d8800b7

Observation deec60cd-dbb6-4a14-8959-22d97b844292 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.142677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.142677Z digest=sha256:aab4a911cb0026b338622cd3cea36445c15f03a08870841272a19aa69b28dfb0

Observation 1b8a3d81-e2c2-4880-a60d-ad3efc81c50f · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.161173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.161173Z digest=sha256:a7039c1de686b829315ee9ee41b57154f387532e145e228ab30128b7604bcece

Observation b887cf36-7211-4919-810c-389fa2787d5d · outbound

This paper cites B leu: a Method for Automatic Evaluation of Machine Translation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility B leu: a Method for Automatic Evaluation of Machine Translation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.174835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.174835Z digest=sha256:0ae0ce0e83921682e4d58375f52e7da65ece8212a64a67855962072ae3c67cd8

Observation 3f1b430f-64f8-4a92-a944-e457a4588547 · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.188195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.188195Z digest=sha256:f19e4cc025ba440feef1b01d63af0665ecc6c5026a73be2a9e2395f7df301bc5

Observation b723fd48-a3b5-4b29-9a75-ec7392e48ea7 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.205102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.205102Z digest=sha256:bb31acd0c76b2eb0e16b1b9f87e390a13632838598c65dff2078e8577a97b016

Observation e0020c01-3cba-4f4a-9b7c-296a91bfd21c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.240825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.240825Z digest=sha256:6a42d9a3f91f6104eaf367146daa5d42d6ac9d08f5a39e71aabbb31222d0351b

Observation 3863b5a4-b9cc-40b4-abc3-35f5ece7cf82 · outbound

This paper cites Advances in neural information processing systems , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in neural information processing systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.253273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.253273Z digest=sha256:09f0e8a0e18bdab7d9464462640380cc5a90cf07fbcb6d1d52759acf327e4fcc

Observation b0d85b71-c52f-4c4b-883f-78f86182d36f · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.265846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.265846Z digest=sha256:560dc42c1c235dc3e5a4d97cf2995e38639a5359672ab37e73cb80a05de2e055

Observation b83f2adc-29e9-45d2-87ac-e8ed9c14eda2 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.276331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.276331Z digest=sha256:9d6fd7eff68d3969c909fd23587bfe1a9c9f821b5f3ed6bb3b8c51aaba6b90da

Observation 4fce2481-a609-4069-9706-5fbf63403382 · outbound

This paper cites Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.288313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.288313Z digest=sha256:571f8db9253928bb21396a3afd2d466a5e3ff14b14bdf6fc98bddc8eb67164be

Observation d741da6b-030a-48f8-a57b-f465dc501039 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.301220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.301220Z digest=sha256:d2946c8e37b2c901c118cc7e37601dffc37d150b9f6fd16dde36a660eb8030c8

Observation bcd421e1-7f19-4c41-a407-3ff5fa179cb6 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Fourteenth International Conference on Learning Representations , year=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.316125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.316125Z digest=sha256:d1cab3f3d7735e2175b682c141c5e9d1e2791a8ba6513b067dc4248b561ed7c2

Observation 535e6139-6e96-4916-b79b-def2626253b9 · outbound

This paper cites The Annals of Statistics , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Annals of Statistics , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.341311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.341311Z digest=sha256:dc508faee30bde774b54c6d5bbcf72fefb9319901c0801a517b2ce62896bd389

Observation 6a62d4e9-d075-4824-93f2-a765cf32bd0a · outbound

This paper cites Forty-first International Conference on Machine Learning , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Forty-first International Conference on Machine Learning , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.353746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.353746Z digest=sha256:c210eb2486d8102fbe744fab9c8c043504f15a15d794d52948e511c933bb622a

Observation d044c412-1ca5-406c-b8dc-a71f4dfdd7ef · outbound

This paper cites A Statistical Framework for Ranking LLM-Based Chatbots.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility A Statistical Framework for Ranking LLM-Based Chatbots

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.364792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.364792Z digest=sha256:6b1d20e8614196bb2898e22e54f8f5d7b68dae5b28790868e925c55d28e42040

Observation c035fa7d-3389-46df-8c86-5a2f9fbb9662 · outbound

This paper cites Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.377159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.377159Z digest=sha256:f88c12b9385934bdf804fe29ec9ba9af90ef133dd50d95f9ed39327aa009dc2d

Observation deb6767e-a044-4401-b2cf-9598a2fce834 · outbound

This paper cites 2nd Workshop on Models of Human Feedback for AI Alignment , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2nd Workshop on Models of Human Feedback for AI Alignment , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.398215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.398215Z digest=sha256:a1b326740953d5fcf712c2aa201ace616b2cd2b55b9981f5ad78a6634623e0fd

Observation be1aded2-6283-4387-927a-e5c1f6e35417 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.417664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.417664Z digest=sha256:079dd2c8dae33a1910eaff41f67a2339c95f5272849f51a152c3fba05c6420a4

Observation aeb66974-4ffe-4bb5-818c-cc2bb35510fe · outbound

This paper cites Improving Reproducibility in Machine Learning Research (A Report from the.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Improving Reproducibility in Machine Learning Research (A Report from the

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.426647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.426647Z digest=sha256:ca5a90fc5d051d321f3e0bcedad0daf0695be54a53bc22546731338d96c4c666

Observation ad7954dd-7d55-4389-ba08-db58a481ca6e · outbound

This paper cites Proceedings of Machine Learning and Systems , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of Machine Learning and Systems , volume =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.448648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.448648Z digest=sha256:0f6afd2ccfc07ed6dda80b8bcd91b5daa18710d9b7c82ce693d705b3ea561827

Observation 7b8f2e42-ac33-4335-86d9-61c10e30b6d9 · outbound

This paper cites Second Conference on Language Modeling , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Second Conference on Language Modeling , year =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.468322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.468322Z digest=sha256:c0ce3625baa891ac0f98d195d38de44b3b5e368b13bd7c2480ebae45cc3ee2d2

Observation fb471102-fd09-4971-ba7f-609d09589d2d · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.476117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.476117Z digest=sha256:055315c4791348500253f878d7c46a06c840948f0db9520e87a4548e1e958162

Observation 21dfafae-aaf9-41de-9244-23aa91f4628d · outbound

This paper cites Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.483178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.483178Z digest=sha256:a2d5fca58174602689ff03fbf6a99cc3ea8b7a6be5e79992c56841b770e6aa99

Observation e4f141b8-ff04-4dae-b1bb-35acbed02698 · outbound

This paper cites , booktitle=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility , booktitle=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.502063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.502063Z digest=sha256:3f2861ee3ca2be98e61b53be7f6ac718f37e8fb7c267bbb65a1f5a2fe276be3c

Observation 15c95db8-f6c4-45c2-aabe-cdeae021d0a0 · outbound

This paper cites 2025 , archivePrefix=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2025 , archivePrefix=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.508897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.508897Z digest=sha256:203737fb45215f759985f37d009ab1463cc9aa4eba4cd7a964ba4563269b4735

Observation 744ab445-8875-4d6c-af4f-b3b486e8406c · outbound

This paper cites Transactions on Machine Learning Research , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Transactions on Machine Learning Research , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.518460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.518460Z digest=sha256:fc7337ea4d9e88faf535bd441b7bd193a21f806af11a8d1569108415f4108091

Observation febafdcc-4389-4721-8159-ec2b15fd27bc · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.530899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.530899Z digest=sha256:02d5003bfaf7bca226297092aa490655181d44a453a27446e1ef290d9d6311cb

Observation 2da06781-b7d6-4afb-bed0-80fbdb395e4a · outbound

This paper cites International Conference on Learning Representations , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility International Conference on Learning Representations , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.538941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.538941Z digest=sha256:73875143d7a67f6e7e92ffd7ab5413c35442135c9e73b8394cddcf789e60b120

Observation dfabeeec-4c98-4079-8907-19d85a39c81e · outbound

This paper cites Nature , volume=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Nature , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.557934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.557934Z digest=sha256:68f3203d18c59f84f641467199b27e442f4398f3f3d33f1e0c3cb265a6e81df8

Observation 859ed12d-2a8c-4e0f-95e8-38b563a9cc09 · outbound

This paper cites International Conference on Machine Learning , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility International Conference on Machine Learning , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.571013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.571013Z digest=sha256:f9e16e541fff32814799e2cdd239af7733838c1759c9c5aa32ce891bd2366d73

Observation 140c94f3-3238-4563-8827-73dfb990a5a2 · outbound

This paper cites International Conference on Learning Representations , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility International Conference on Learning Representations , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.583744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.583744Z digest=sha256:b9a4f6ede9443ba5d759772602de365111ec59d785a91930e84bed79c711a54b

Observation 5e5f0879-9b08-4e83-8dca-8a90ad9890d3 · outbound

This paper cites First Conference on Language Modeling , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility First Conference on Language Modeling , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.594268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.594268Z digest=sha256:fc6585d0559c415ac3bc1b7fe39f09a00f042cab7f9e40ba7f4a727ef7c4a6b1

Observation aafa9cc9-6c4f-4924-80b9-bbf6fd078530 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Instruction-Following Evaluation for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.604718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.604718Z digest=sha256:2f2ca4baa24f2cc2bf6bc6a8cec7fc2484fae87807436772b6eed4812a3905ab

Observation 4ae6fece-e2b3-4f0c-b451-ee0c051d3c09 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.620654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.620654Z digest=sha256:aaa2b5976792a47bf3ba2ea912f35dff41731b676381d7d44947298588dcc351

Observation b91102af-3bb0-4ed8-8aa6-b13203c9c064 · outbound

This paper cites MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.637202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.637202Z digest=sha256:34bf659a5a793e2ef18bb41a1f6770344f0d50b8e169a89cf6efaef6eef9b495

Observation 465eeb61-4db3-405f-9654-1203bd1759fc · outbound

This paper cites International Conference on Learning Representations , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility International Conference on Learning Representations , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.650955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.650955Z digest=sha256:e2fb697b2b9c88df8447259bdc1d29c3c8a460a183a407fd7adfc7e98fb44d31

Observation a43964eb-31bf-490f-94d5-0a9d68e2461d · outbound

This paper cites arXiv preprint arXiv:2602.10367 , year=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility arXiv preprint arXiv:2602.10367 , year=

Reference 51

Resolution
verified exact
doi, observed 2026-08-15T14:49:10.171236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:49:03.665804Z digest=sha256:d5d463d337b74663a55b1e204d4e6165269846071e5c5c2e81f5e95a1fda65d4

Observation cf1eea6e-0a93-4652-a4a6-1969a1400f5e · outbound

This paper cites Challenging.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Challenging

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.677534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.677534Z digest=sha256:93a2fc9bba76227b2f9a57d2b3e16d0524cd73227e9d9050246979957ff53e2e

Observation 187f9a1b-360c-4700-9c77-76e077d89bc5 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Measuring Mathematical Problem Solving With the MATH Dataset

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.690322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.690322Z digest=sha256:7049d7855b3042b73c644e56f78a21a79ad24b611b8441803596a382223e6781

Observation 7e4a299d-4bec-466d-ab03-e131870d2d61 · outbound

This paper cites 2025 , archivePrefix=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2025 , archivePrefix=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.705214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.705214Z digest=sha256:36a78ad12e202d49787279e0a92efe8dbdffe7f3a298469da48357d34ccd2396

Observation b7df270b-d100-489a-8eec-9c32b809d689 · outbound

This paper cites 2026 , archivePrefix=.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2026 , archivePrefix=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.714669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.714669Z digest=sha256:7652548c15e4fa90b86347eb9d3536488cd68279ae40cb8703ac50a0e0ead622

Observation ca48349d-395d-4946-98d5-122b2d80fe5b · outbound

This paper cites 2026 , doi =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2026 , doi =

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.722278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.722278Z digest=sha256:97bb5435f98e166f8b68987158db235776a7ebf7ed0fb7dfb26902d9459a4f09

Observation 7b8a6c81-ce65-4b6e-a8b4-fddee27edba2 · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.748051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.748051Z digest=sha256:72ea00c20f23a70f11449e0ac32d256c317893b077e6ed04f29599739d422b86

Observation 31dc3ccd-c9b6-4ff9-ba90-214a4f4677ab · outbound

This paper cites Qwen2.5 Technical Report.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Qwen2.5 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.759711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.759711Z digest=sha256:e18c9041ba01b8197d1a61df890689ba8f5334ed2dc2686906143b5acd7d5287

Observation 8afa8fae-5397-41e3-b0b5-aac3e2051f8c · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.772795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.772795Z digest=sha256:6328257025de7de2bfd53f4f9e52c7ea443d042a84768060c6c9ec1807229344

Observation bb3c1410-5ba5-4936-98cd-b4636762f556 · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.789033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.789033Z digest=sha256:3576774af1a81629b39ba1e31712ed9ad5cf96dedb853819cea0465e4c9a7987

Observation a6195a13-6e25-4577-b6c2-23e65d0920ae · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.810806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.810806Z digest=sha256:911e7903007e487e1612f1576c25879be35a1eeea2475be029f085628918507f

Observation 9897b4f5-f0cd-4f79-8b44-358fd6a9cffa · outbound

This paper cites S*: Test Time Scaling for Code Generation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility S*: Test Time Scaling for Code Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.822238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.822238Z digest=sha256:efffcf0fea524de7209cbb04f6bbfd6fc19027418b4bfea19ce0af1fc253e222

Observation 9467578a-49db-46af-970a-74d3768ec7a7 · outbound

This paper cites LIMA: Less Is More for Alignment.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility LIMA: Less Is More for Alignment

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.850897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.850897Z digest=sha256:e31744253617afc0d1c93158a3a87b195047a389cec8f38106b4d667151c1e47

Observation a8024bfb-8703-4df7-be90-c56f4140dd78 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility LIMR: Less is More for RL Scaling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.884561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.884561Z digest=sha256:2e27f0f513daf0d65da9a04fe81813525cfe31dc54bf66ea685afff621ddbd07

Observation bf26976d-e44e-426b-a78b-5bcac61826df · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.892518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.892518Z digest=sha256:3473e546f87818ce98b905dc504a61dbbbf6f5bcc0adb7e36a2ccc1860391416

Observation 2aeb59e9-da0e-4608-9a67-7518b7e2697b · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.900136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.900136Z digest=sha256:a17cf93c32cae4567fbde2784027c687074ac96a9425affcc883633e8feb5ff6

Observation 380d3236-b9ad-4d63-933f-42094778f7fd · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.912354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.912354Z digest=sha256:cec534062bc24661766869ee2615aea370a55e0dcf28ce233638e0c2517e9927

Observation e023607b-b724-4b2f-93c9-1862ccaa2d65 · outbound

This paper cites Knowledge Fusion of Large Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Knowledge Fusion of Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.926672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.926672Z digest=sha256:eee4d1ccd693a2bcdf269a336ec8f025bde0fd3e02c683d88abe014b405f0def

Observation 332bdf63-fe15-4976-9ed2-77e40cc2b11d · outbound

This paper cites FuseChat: Knowledge Fusion of Chat Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility FuseChat: Knowledge Fusion of Chat Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.942405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.942405Z digest=sha256:7f964e65b763511e8a5e5096e233a07ca2a3849aa70de6a694cbb33106e3b0e8

Observation a3bd99cc-ceeb-449e-981c-92e22641e042 · outbound

This paper cites FuseChat-3.0: Preference Optimization Meets Heterogeneous Model Fusion.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility FuseChat-3.0: Preference Optimization Meets Heterogeneous Model Fusion

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.966561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.966561Z digest=sha256:cd11f5df15c3acd40ebb99938f2d7eba7dab094a356ecc127b21e70dbfc43325

Observation 1eafb797-f179-47da-97b1-418540031a12 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.983309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.983309Z digest=sha256:401546d9915636c25d2086a370ff575448c3b93070a88b171513f25eed014621

Observation df092afa-de0a-4486-af45-a42e10b4aa6a · outbound

This paper cites TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.000351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.000351Z digest=sha256:a36977357c28afab0b86b9ccd7d41fb7aebce41f67050d620ece5495e5759b35

Observation 809b3a86-bf6f-4c0f-9fce-ae93cedd94bc · outbound

This paper cites The Twelfth International Conference on Learning Representations , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Twelfth International Conference on Learning Representations , year =

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.026475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.026475Z digest=sha256:238651c1beab2ef8bf7543978ce91227ed0f4cc817cd13f764daca24a3f68c78

Observation 478271a6-d874-498f-87f3-4be5a1de9795 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Fourteenth International Conference on Learning Representations , year =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.036564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.036564Z digest=sha256:f40a8e3b0a89812fd7f19da0abd7735b5b7af28e202034489b7f71fcc513b354

Observation 2ce0b4eb-d728-45e2-9ffb-879216c3baa0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume =

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.052419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.052419Z digest=sha256:93c16acf2c78145fde18d99a9745cc068edd163613c53fb8de82bbe552e92410

Observation fdfb8f6c-12eb-4cea-91eb-3be546c48d71 · outbound

This paper cites Optimal Aggregation of.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Optimal Aggregation of

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.065643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.065643Z digest=sha256:56399019e99933ee6ef88ba44b0136c529c211cfbce80f4a378ac1d8488cbcd5

Observation dacdbd11-2b29-464c-8d42-666e13e8adb3 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , volume =

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.073304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.073304Z digest=sha256:1cf10e3dad798791afc52dbce5211b4c640a26299fb01a4499514561d0bab0cf

Observation 125c9ea7-6e1c-4a21-ae1f-7eb909cd7fd2 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility The Fourteenth International Conference on Learning Representations , year =

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.084816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.084816Z digest=sha256:1228f6bc61f74a17bd89ee2917c1be06e30845868d12336d1e350eef9ac053f0

Observation 77f8467b-9cc6-4d44-b7ca-fd51bec82f1e · outbound

This paper cites Transactions on Machine Learning Research , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Transactions on Machine Learning Research , year =

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.100667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.100667Z digest=sha256:a3af386f3ac11210b7675db3cb4c25117b51863183db9b742f456780bc67c12f

Observation 71f7c46e-1a9b-40fe-906c-5556af1ea1ae · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning (ICML) , series =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the 40th International Conference on Machine Learning (ICML) , series =

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.110984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.110984Z digest=sha256:a26d23d0e6acbb519f280127b3fbf876b9586e5b6fe1e548474e3d9adcfdf0ad

Observation 5387bdf4-bcf9-4feb-a438-ad11c386f66b · outbound

This paper cites Proceedings of the 42nd International Conference on Machine Learning (ICML) , series =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the 42nd International Conference on Machine Learning (ICML) , series =

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.136460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.136460Z digest=sha256:1749507ee50d7dbfe16de9d8862f7c74926296b2bbfe8a68a5d1e937d3be0b43

Observation c56500bb-cc81-4324-9519-2830e00b443c · outbound

This paper cites Soft Best-of-n Sampling for Model Alignment.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Soft Best-of-n Sampling for Model Alignment

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.163654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.163654Z digest=sha256:2151bd2226da9131e9e437c4e987e46158c1ddf3f97fdd4b77313f732708c6c9

Observation 92d5ee9e-8602-4389-b1ed-c0cb7aed0ca3 · outbound

This paper cites It's MBR All the Way Down: Modern Generation Techniques Through the Lens of Minimum Bayes Risk.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility It's MBR All the Way Down: Modern Generation Techniques Through the Lens of Minimum Bayes Risk

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:49:09.925826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:49:04.177578Z digest=sha256:c572a7e6d96368fa03793b2245effa0f01ee1111507716cfd554695905d16eca

Observation c673f3cc-00dc-4e2c-9e02-904c41bb1739 · outbound

This paper cites Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL) , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL) , year =

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.188855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.188855Z digest=sha256:affe909ab1fc3009a2f0ffb9a52bd3ff9855edb00305f48ea994e939cabea8c6

Observation d436f7b9-b20f-4fee-9250-0b5715addc7e · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Transactions of the Association for Computational Linguistics , volume =

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.197812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.197812Z digest=sha256:c2189388c26d71bc35ae371a2cc4cce855adb9fa6e80ee41d458e569e07f93f3

Observation 01295ef8-cac2-4f6d-9a0f-264133eb254b · outbound

This paper cites Better Instruction-Following Through Minimum Bayes Risk.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Better Instruction-Following Through Minimum Bayes Risk

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.204164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.204164Z digest=sha256:7dfeb0473e997dab89a7f350a8d29a985844b246512ac346543e929052ae8986

Observation 07c3911c-76bf-493e-8842-6f83b1d2d69e · outbound

This paper cites Faster Minimum Bayes Risk Decoding with Confidence-based Pruning.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Faster Minimum Bayes Risk Decoding with Confidence-based Pruning

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:49:09.780154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:49:04.212031Z digest=sha256:38e27343802bb5cd44fc070a7708402610d554ba02950f867330529b125ddf52

Observation ef98f7cc-dce0-4a33-82d8-643987a11e78 · outbound

This paper cites Hyperparameter-Free Approach for Faster Minimum Bayes Risk Decoding.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Hyperparameter-Free Approach for Faster Minimum Bayes Risk Decoding

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:49:09.689468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:49:04.223920Z digest=sha256:a6d194354530475d228acff4daa44a46db0363afae8533a767e3435d6e3c0a41

Observation e659d0c6-c363-4598-b954-9181caf8ad6f · outbound

This paper cites an unresolved cited work.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.267671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.267671Z digest=sha256:099737fc5879bf981caf82c9f3e76c046defaa2c59955b8e82b77ce5f93fcc86

Observation 6f644802-bea2-40f0-b29c-f90a475d5cf4 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.283055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.283055Z digest=sha256:227a4f1db29a915d1c7c8c514c8bdf88ebc3a8e1cbfc50ba992efff0a7e8ec52

Observation 04540a5a-9b77-48e7-973a-0bae09f5332b · outbound

This paper cites Universal Self-Consistency for Large Language Model Generation.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Universal Self-Consistency for Large Language Model Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.294789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.294789Z digest=sha256:9c360aa4556071ad08b01a8ff2bae4a604c2a338265c80d52d012fbe737d76ec

Observation 5203ef2c-225a-4bd0-83bd-a693844b4bb2 · outbound

This paper cites Ranked Voting based Self-Consistency of Large Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Ranked Voting based Self-Consistency of Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.410743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.410743Z digest=sha256:734d3adf5c99a40ae817ac508165abb0b5262679a82f322b9f49c609206040b6

Observation 4bf09981-9f9c-44a9-8920-c77510fc6931 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Training Verifiers to Solve Math Word Problems

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.448798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.448798Z digest=sha256:5313876dc40f6c073d599bd9022e4a9a3649e38b94022d99901bb1334480ad6b

Observation 982ad22c-feed-4b18-939b-2f880fec4c88 · outbound

This paper cites Let's Verify Step by Step.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Let's Verify Step by Step

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.469228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.469228Z digest=sha256:660f376058a36f231b7ed2e3ae8e14d9d6893ab6d6cadf29402a1f31c64e050f

Observation 61a36d49-167c-4402-8516-20f0e84765a4 · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.482328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.482328Z digest=sha256:828ca6d62e93bbe6853c502588e391d993ab8779cb7ff25ec735549245c1ae4f

Observation 18fdb6ed-7ca2-453d-a9a1-399b5c03c232 · outbound

This paper cites Variational Best-of-N Alignment.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Variational Best-of-N Alignment

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.493890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.493890Z digest=sha256:5c996341039b91481e58b7387f4b3076fb9c260ca56bda658cbc9eb2c6a0fa5c

Observation 9c4f0e19-75d9-4289-b4e4-913bfb9ff9c3 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.507240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.507240Z digest=sha256:35b4da29e6011ed6a792a2ce69478346d76110451064d1384984c1629117bb06

Observation de297c5c-2d83-4dc2-a45f-fdb00483a7e6 · outbound

This paper cites 2024 , month = jul, publisher =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2024 , month = jul, publisher =

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.514114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.514114Z digest=sha256:778307f5da0b63523b0cf5babdc01204ca4dd5be668c497a963df179f92535e0

Observation 7b9ed27f-c2d5-49e1-886f-7968f8de9626 · outbound

This paper cites 2024 , month = jun, publisher =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2024 , month = jun, publisher =

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.520452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.520452Z digest=sha256:7293253df8bf8df6f62df29a5b986786f10e129f4622cc31f0f9961bb88ba467

Observation ea09fda9-5d18-4c4a-b8e4-1fff5b46108a · outbound

This paper cites Policy Guided Tree Search for Enhanced.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Policy Guided Tree Search for Enhanced

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.528030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.528030Z digest=sha256:bfb008a4d4b722126337bfb84d31fddbbb8f23efd650c50dcedfed01c8bf89c6

Observation 88595a3f-38b0-42c4-85c3-636c7d31c86f · outbound

This paper cites Transactions on Machine Learning Research , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Transactions on Machine Learning Research , year =

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.536806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.536806Z digest=sha256:aa02a5b244c716b1b0d3cbc693852510e3718aee9dff36cedfd46d294f278bce

Observation e987dfd2-6f6a-4940-8dec-f459990e9a28 · outbound

This paper cites 2026 , eprint =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2026 , eprint =

Reference 107

Resolution
verified exact
doi, observed 2026-08-15T14:49:09.041995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:49:04.543603Z digest=sha256:8d37b048fecea499342e3dea90117b8fc16d160fb53bbc015e43d2f92592e3f4

Observation 6b0eb9fe-4132-4c92-9fd7-a1b39b19e39e · outbound

This paper cites 2026 , note =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 2026 , note =

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.550606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.550606Z digest=sha256:d531e661af2719694d60e7b48ae2298af1c90c7cee020538709fd7566ca2d04a

Observation 70cacda5-6998-49f4-ba8a-15bb4988c998 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , year =

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.558411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.558411Z digest=sha256:9da8aba32cfb95c7a657c7a8925fb49349aa1557a00471c227bb1e206481a581

Observation 28e6f3fe-ca38-4975-81f4-a9f712a87b4a · outbound

This paper cites DEFT: Decoding with Flash Tree-Attention for Efficient Tree-Structured.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility DEFT: Decoding with Flash Tree-Attention for Efficient Tree-Structured

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.567845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.567845Z digest=sha256:a55990417d6391f4a5c98df7f3bd1e05cbab317d5dd370df688c2512bce9dc42

Observation 73740463-84e4-4c4b-9342-7730f1b5dce3 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Advances in Neural Information Processing Systems , year =

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.576930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.576930Z digest=sha256:9244f46e38f00e50c5c7c3cb7a9a53a128a601acae11a2670faafe394081b247

Observation 021f8166-9156-4f2d-a5a3-4a5afde8ccc2 · outbound

This paper cites Bandit Based.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Bandit Based

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.587472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.587472Z digest=sha256:20e01c45228dcc878e782385ad6710148fa71de3d1470ad85525d6e588ce4226

Observation a2bcafe9-21f7-4526-bda4-f0ce31176a44 · outbound

This paper cites and Powley, Edward and Whitehouse, Daniel and Lucas, Simon M.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility and Powley, Edward and Whitehouse, Daniel and Lucas, Simon M

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:04.602845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:04.602845Z digest=sha256:08da57b78b644ddbc8f153955f7f54b2596dcc6beeac352c1c54eba0a068df5a

Pith citing papers

No inbound Pith citation observations are available.