Pith. sign in

Paper Citation Record · LEDGER

Reward Model Generalization for Compute-Aware Test-Time Reasoning

As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2505.18065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18065 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:27.675920Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d07cc6f-051c-4d59-b313-62ffd2c6aba3 · outbound

This paper cites Le, and Denny Zhou.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, and Denny Zhou

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:33.454821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.081236Z digest=sha256:8443bef624a13f15ce8e97fbd8463d1906cdb3041f4a767f1d26953144645de6

Observation 05fdb4ef-c227-4761-88b5-7b958e0d5059 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:24.152490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:24.152490Z digest=sha256:427e2176ad732dffcba8e96f4ae0d2fa25ab883379ee83ccbaa2a14f1ed84d6e

Observation 727b13c4-1c25-4be9-9086-42328e456b6d · outbound

This paper cites Griffiths, Yuan Cao, and Karthik R.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Griffiths, Yuan Cao, and Karthik R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:33.163810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.250753Z digest=sha256:3f16beaaee8641c69f52c1435010c88485f836b1242f55671ae3f638aeee2cce

Observation 18a474a5-bc81-4dc5-9c91-06cbda94bfb9 · outbound

This paper cites Learning to reason with llms.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Learning to reason with llms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.923839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.362104Z digest=sha256:ab6a53128414c8686639daba09424d0288c34480cd61c2f50d3acfe0a08d7af3

Observation c3a4d3ef-8136-4f88-b5b5-ae8006654fbf · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, January 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, January 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:24.426613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:24.426613Z digest=sha256:b16fa82295d4e383418495e43649781b9a41a640b25f6c413df4f5bfdc93a848

Observation 13428347-59ee-4da7-aba4-b08a2574e3ff · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:32.681900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.534100Z digest=sha256:90f3acd5f7b7a7b3fe7e4729f115f0b1e08a8aa8b92f5037c135ca4f15e0a684

Observation a0dfc639-8053-4ea3-b46b-f059cd7739c1 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.593348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.600772Z digest=sha256:1ae324319b02b1bdbac731c770a9383df3e40606561e295b32e38d2a43685a9e

Observation 6ecda57c-0ee5-49bb-b990-d2760e9246bb · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling, February 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling, February 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.491605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.687363Z digest=sha256:4c441d991844acc4f2d32d4e5fee9363068d0d6971ecdcb74230594ac735b75b

Observation fc930b81-756e-43f1-ab68-5a3c5057776d · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.372697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.806743Z digest=sha256:4899ad2bc49f9283dd1e9c020fa7be11a8464447a90fda1b3ce36b3ec12feec9

Observation 8d6b8157-deac-40a5-9ae5-9e66c590011a · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Graph of thoughts: Solving elaborate problems with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:31.491532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.911167Z digest=sha256:fbb8226324517c276ccfd4bcc098b550e38f290f98d45bd775c1a36d735b0e44

Observation 22167700-eacc-421b-b6dd-ad8d092ee810 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:30.980423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:24.986744Z digest=sha256:b8663fda5077aa52bc826d2848bbeb733acb98d382080e2edfdf4bb0576955f2

Observation 4d3043e1-1d95-4555-9bfe-6683a1065dcb · outbound

This paper cites Let’s Verify Step by Step, May 2023.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Let’s Verify Step by Step, May 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.665499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:25.116168Z digest=sha256:6b2d667e7214c400846d7272c13a4335a80ceaedfe1911e64a2ba71b683d1fa2

Observation c38bed93-daca-4fd4-b267-9bb599ca8ab4 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:30.551034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:25.192718Z digest=sha256:0728f1b244a65a3e7404939677f4d7100470914dcf9d1f7f43535ce2ade6d952

Observation 5a24e95c-eec4-45e1-bb29-4a6f6118c9f5 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Measuring mathematical problem solving with the math dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.297848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.297848Z digest=sha256:5ba04231093b6abb84e60fc501eecff70a3100e69d7602140f1cbceb6e3e60a2

Observation 0f2507de-19be-4ef7-bfb3-b53c6b5ec71e · outbound

This paper cites AIME 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning AIME 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.371400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:25.365975Z digest=sha256:727d1d7e94b80e059e809f6c27b811f17b6bcd215ae6c36c998dcb66abc9d815

Observation ed7f689d-b065-4e0f-b891-59a0eedcc988 · outbound

This paper cites Qwen2.5 Technical Report.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.458387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.458387Z digest=sha256:ec5242360b9cf6999c5ef95cf9bd63cb2077550beb1db38e3defa995a8722e1a

Observation 12682f0f-c37d-4d95-9637-82eaa88d51f1 · outbound

This paper cites The Llama 3 Herd of Models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.571342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.571342Z digest=sha256:86a103d019f8a6758bb886c57f29d18f0783ecec2ab2455dca1705bb2d09964a

Observation 97dcf570-6c88-4f66-abed-e1c38a9228d1 · outbound

This paper cites Llama 3 to connect 2024: Vision, edge, and mobile devices.https://ai.meta.com/ blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Llama 3 to connect 2024: Vision, edge, and mobile devices.https://ai.meta.com/ blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.257185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:25.638982Z digest=sha256:c602b244995dcbf86709a2100bb468abe5f8884a99139079ea60ab5b22ba393e

Observation 111de10c-420a-4d6e-bf88-52694d5cc6c8 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning, January 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning, January 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.155577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:25.709602Z digest=sha256:c0fd1f456be9e6cb2a991977680c043f5dfcbf39f5aa82cbd63a0dc3fce8b5ba

Observation 5bfd9ae4-d6fb-48ab-9074-776f40b33358 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Reward Model Generalization for Compute-Aware Test-Time Reasoning RLHF Workflow: From Reward Modeling to Online RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.802972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.802972Z digest=sha256:a313403b1473908b40c8fb53518a7bf351ebd7c3ae8c23de293c558b474b6447

Observation 9b4d4216-0c01-49c7-bc2e-deb80b3ec527 · outbound

This paper cites Skywork-o1 open series.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Skywork-o1 open series

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.991155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:25.912177Z digest=sha256:b25446061ae9e843bc757db4208786abb30c94bc910be345391c09b62228e24e

Observation 63e3dcf3-98ff-48b8-b892-fd046de0b922 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models, April 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models, April 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.885755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.045623Z digest=sha256:d4d3ab8058eee3d429becb69b0ce821d3e1fc31fe178ca752a28c0916fdf5f6a

Observation b298c93a-8765-4701-b57a-9268218fcdb4 · outbound

This paper cites Self- Refine: Iterative Refinement with Self-Feedback.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self- Refine: Iterative Refinement with Self-Feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.779220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.134417Z digest=sha256:f29229d961e54f4f2f5b2c8620fb78e3391f872f77942bf09d07fcd12aa0f892

Observation 79eae5de-7d27-441b-9245-a15d3ef2d0c8 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self-critiquing models for assisting human evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.240819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.240819Z digest=sha256:cf36ff7ee74ec4f1792afb33895f749fd8d9f0bed391322b55dea9f24827fda6

Observation 4ab4ba77-f6a9-464a-a491-97db88620f98 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:29.651910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.313976Z digest=sha256:daa4aedf7a16b7ba7ff2bbe2a175ca8f890d539dc80078eb7d061eef4a1ea2da

Observation d1580805-140a-400e-af05-3e50ae1c20c2 · outbound

This paper cites Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.538636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.414898Z digest=sha256:868c11520e8098576fafe3b00331017300175f2b2536d9b78cae2439994b85a6

Observation 23457b43-4bf6-4ab9-986a-95249c5d837d · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models, October 2022.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Automatic Chain of Thought Prompting in Large Language Models, October 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.403676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.484880Z digest=sha256:f765c065487a5a995c7c3212fcb862f8a6fbf097a118403f1fefa7bc2d0da25c

Observation 3e53c68c-5b28-483c-8688-904cf476686a · outbound

This paper cites Le, Christopher Ré, and Azalia Mirhoseini.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, Christopher Ré, and Azalia Mirhoseini

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.235347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.586366Z digest=sha256:282683a6292d464282123d5e2fd804435e6c80049b9debf6337837a59a37e598

Observation 80aeda13-4e57-474a-830d-9b7989e22c95 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Solving math word problems with process- and outcome-based feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.655098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.655098Z digest=sha256:5de0324c9443c1704837f61b7c6c8c654aff4fdf3d1f777db2217d92fff8d955

Observation 3f49e629-983f-402e-8de3-c6bd1a6974a4 · outbound

This paper cites Self-Evaluation Guided Beam Search for Reasoning, October 2023.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self-Evaluation Guided Beam Search for Reasoning, October 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.066930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.752346Z digest=sha256:5b66c31126c220ccdae70c22cf066f0d46a87eb9bbd93c24e1c3413a49d65891

Observation 5989dbf9-6351-427a-96f7-a2b83417e589 · outbound

This paper cites Le, Ed H.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, Ed H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.888280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.821415Z digest=sha256:3c1b8cd2c1facfe41b2446e91bd6c58752cb165ec0eb22b57258f3edcd496e58

Observation 0f63f736-7ba2-4871-bd45-e37bdb4fc507 · outbound

This paper cites Sparsity-aware generalization theory for deep neural networks.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Sparsity-aware generalization theory for deep neural networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.710906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.878196Z digest=sha256:c1fb8ce54b698678021438d77041bdd7e1271940ffb04060b511a744ed105990

Observation bcd5bd05-065f-458f-8156-47c1acb6cf61 · outbound

This paper cites Pac-bayes compression bounds so tight that they can explain generalization.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayes compression bounds so tight that they can explain generalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.529245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:26.941110Z digest=sha256:a1e3ff1e8a777b7c1c637bad7d1f01d94bad5d7b06ee8dee2ee9f9e772320ef6

Observation 4310f865-1e8a-40e8-8302-6476b744c6a7 · outbound

This paper cites Efficient content-based sparse attention with routing transformers.Transactions of the Association for Computational Linguistics, 9:53–68, 2021.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Efficient content-based sparse attention with routing transformers.Transactions of the Association for Computational Linguistics, 9:53–68, 2021

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.981219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.981219Z digest=sha256:3b0d2fbbcfa7cbc533038f509155208e04def1cc087f3bc92c2c920bc771b4e7

Observation aeec097a-a6a7-4685-b0a1-af1c8d334661 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Reward Model Generalization for Compute-Aware Test-Time Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.030634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.030634Z digest=sha256:a5c4ac5fd601b908c4e97fd55284e76175e1415150af6c2065ce2d3da25b2af9

Observation ae4a9032-b346-4db8-893c-4db624af8ef8 · outbound

This paper cites Policy gradient meth- ods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Policy gradient meth- ods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.086932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.086932Z digest=sha256:87fdfea4cf22807712f05a0eab7dbb03fae3781e34122b100bdcb6bc6a9a0ee2

Observation 8ab56fe0-8d44-4325-bf39-5dde9dc87c86 · outbound

This paper cites Pac-bayesian model averaging.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayesian model averaging

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.135988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.135988Z digest=sha256:b3f95a1cfb750c802024628b4b9955bb2258c6b7e595bedbf0ec828a9aed16e5

Observation 908e7d2b-f6af-4cf3-ae26-283ddea24f6a · outbound

This paper cites Pac-bayesian generalisation error bounds for gaussian process classification.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayesian generalisation error bounds for gaussian process classification

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.359430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:27.172308Z digest=sha256:3846f2338e654a70c6f70ac086f0aee868066f74b07e5b923a37537f223107bd

Observation bd957e38-9971-47c0-9ac3-95acd361c0e5 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Training Verifiers to Solve Math Word Problems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.222407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.222407Z digest=sha256:d543b498cbf74a627652a3fbb9566d7645b84fd9df26d0889d57fe87ddddb41f

Observation 88c52a6c-370f-4d0c-8894-c2d94f5e85d4 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.263006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.263006Z digest=sha256:561e18dab02a8400006ac201f83a4ad56f3ca1d392f83cdc9c12d16346d12b13

Observation f9b003d7-32fe-4b32-ab05-8497d578fab5 · outbound

This paper cites Enhancing LLM Reasoning with Reward-guided Tree Search.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Enhancing LLM Reasoning with Reward-guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.309575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.309575Z digest=sha256:f7ed2663d942ac3513c08df6f62e762c73841a258e6cf505b54f6863c4097cdd

Observation ff489b9b-8686-4b99-929b-ebd174b217a5 · outbound

This paper cites LiteSearch: Efficacious Tree Search for LLM.

Reward Model Generalization for Compute-Aware Test-Time Reasoning LiteSearch: Efficacious Tree Search for LLM

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.388786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.388786Z digest=sha256:344589102752d87bbf013f6ceda6e8de96ef9f8821fe8e22defb7bb865f4c972

Observation 5deda4ba-41c6-44cd-8c92-812854e08f0d · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Reward Model Generalization for Compute-Aware Test-Time Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.432302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.432302Z digest=sha256:40d6efbf8f638f44e7fd0cb72a60f743444241dd5b6653522eeb79ca75eb7049

Observation b0e8a27f-7ab7-4583-9ee6-bbe4a596520d · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.461851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.461851Z digest=sha256:e11a546cad30a8b8b1a6bff96a2f574f48cf83c294a97ca8068c79649bbbb623

Observation 4ded9f65-0ab6-491b-815f-f82eeb5b2c6c · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.488154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.488154Z digest=sha256:d7d2542af7916deee2f11345af2130f5d307ef666135cd2c97d65a61a4122eb0

Observation b23991e7-dce8-4463-9a82-445e0a635091 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.539519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.539519Z digest=sha256:bf74693fcca3727b68388573c745ace7c9cdafc0d9f09d1d878d2e5c20e98971

Observation 534d5fd5-3165-403d-b888-4f26d9ddf483 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.571108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.571108Z digest=sha256:40894af750139df0b28d6bd9e1c47acc1fa5c660da707426ed90a8fbb7c1ffe2

Observation 55695bc1-c234-4fde-aef1-1f7a950991d9 · outbound

This paper cites BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.601318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.601318Z digest=sha256:371cf1ebc702497a91f1677189745b4ff3e329b74c89fa444591fa9d1a44ece3

Observation da4596c2-6956-4a5a-8812-5bfe8b0fe509 · outbound

This paper cites Subtracting from 1 yields the bound Equation 6.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Subtracting from 1 yields the bound Equation 6

Reference 49

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:42:27.897972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:42:27.675920Z digest=sha256:106966a857d953b48c2c2b3bf5b4f9466d96f03e5396641937b10b634417ad0c

Pith citing papers

No inbound Pith citation observations are available.