Pith. sign in

Paper Citation Record · LEDGER

Reward Model Generalization for Compute-Aware Test-Time Reasoning

As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2505.18065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18065 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:27.675920Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d07cc6f-051c-4d59-b313-62ffd2c6aba3 · outbound

This paper cites Le, and Denny Zhou.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, and Denny Zhou

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:33.454821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.081236Z digest=sha256:ca89e0de2c7c8f9d639fae2266b3d5adb97386bc35b3af9ed0d718112fc70347

Observation 05fdb4ef-c227-4761-88b5-7b958e0d5059 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:24.152490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:24.152490Z digest=sha256:a9a8bc8a4cb215276409a3e63d2b64553dbc20ac6fa28f926c78cedef73a7d65

Observation 727b13c4-1c25-4be9-9086-42328e456b6d · outbound

This paper cites Griffiths, Yuan Cao, and Karthik R.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Griffiths, Yuan Cao, and Karthik R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:33.163810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.250753Z digest=sha256:9458e384151c7fb23d09f7d92a52f7885240030871d66de1ce6f37937c11f225

Observation 18a474a5-bc81-4dc5-9c91-06cbda94bfb9 · outbound

This paper cites Learning to reason with llms.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Learning to reason with llms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.923839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.362104Z digest=sha256:b9b511a26265f77ffe46ef4f7f34cbe31c2275c79317153e537dd0fab1d6d29d

Observation c3a4d3ef-8136-4f88-b5b5-ae8006654fbf · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, January 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, January 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:24.426613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:24.426613Z digest=sha256:c4089e61f101bc8d0d55f19e2403508c1b6546017b57f9f43f46fb33f58a36b7

Observation 13428347-59ee-4da7-aba4-b08a2574e3ff · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:32.681900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.534100Z digest=sha256:e521e4cc9f6a67cbf302f4343c97505f7bca3feb96492a740b689db57df07288

Observation a0dfc639-8053-4ea3-b46b-f059cd7739c1 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.593348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.600772Z digest=sha256:b9c9813033938c2c4a2d5bb3ca13de875b81a41573032a216b948c28c31d56e2

Observation 6ecda57c-0ee5-49bb-b990-d2760e9246bb · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling, February 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling, February 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.491605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.687363Z digest=sha256:42492639ce6cb9ac5e0accc39fd4ac210820931739cbe60476e46c02f8977d26

Observation fc930b81-756e-43f1-ab68-5a3c5057776d · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:32.372697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.806743Z digest=sha256:d88f135fc6faf7986134b0e36ffbb71ebcef8603a8ea9dc2950c417cc11f4e46

Observation 8d6b8157-deac-40a5-9ae5-9e66c590011a · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Graph of thoughts: Solving elaborate problems with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:31.491532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.911167Z digest=sha256:5d30320009a854883b08da38664d840f72bd7ac3ef0b0ef91bb1c2f960844472

Observation 22167700-eacc-421b-b6dd-ad8d092ee810 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:30.980423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:24.986744Z digest=sha256:61d394b23a962a340b4e0c3bfc7e4ae464fd9b783535e570afa45a15f271d1d1

Observation 4d3043e1-1d95-4555-9bfe-6683a1065dcb · outbound

This paper cites Let’s Verify Step by Step, May 2023.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Let’s Verify Step by Step, May 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.665499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:25.116168Z digest=sha256:8485aeca4176199b88ffaba72f5ea0bd33891e1204bd976e3d4371896fb85537

Observation c38bed93-daca-4fd4-b267-9bb599ca8ab4 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:30.551034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:25.192718Z digest=sha256:5d2b5127fa17f6452916cb442794385d9799162b69b689db5d73772db3aaf29a

Observation 5a24e95c-eec4-45e1-bb29-4a6f6118c9f5 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Measuring mathematical problem solving with the math dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.297848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.297848Z digest=sha256:98d83b9a195b17a85246fbca9313fbb5cc98debcfa67a515f422cb047adc81a3

Observation 0f2507de-19be-4ef7-bfb3-b53c6b5ec71e · outbound

This paper cites AIME 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning AIME 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.371400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:25.365975Z digest=sha256:a4167f7716a7bda882743d73f886aa0361ffe17e3ab4e74dba491b92a3680923

Observation ed7f689d-b065-4e0f-b891-59a0eedcc988 · outbound

This paper cites Qwen2.5 Technical Report.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.458387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.458387Z digest=sha256:3ba65486d3d2c07526f5a77a9148bd412952129844a37f58a124dddac2526567

Observation 12682f0f-c37d-4d95-9637-82eaa88d51f1 · outbound

This paper cites The Llama 3 Herd of Models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.571342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.571342Z digest=sha256:d920bf8f3b178a0fabd4004aab4581bab7e9d916161f76e9ff4c0b66e46135ab

Observation 97dcf570-6c88-4f66-abed-e1c38a9228d1 · outbound

This paper cites Llama 3 to connect 2024: Vision, edge, and mobile devices.https://ai.meta.com/ blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Llama 3 to connect 2024: Vision, edge, and mobile devices.https://ai.meta.com/ blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.257185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:25.638982Z digest=sha256:e6bb07495b3c0f38fac31fe8386b7d9ab6ec4d5ffb0d15c6e6ad9dd3ba319437

Observation 111de10c-420a-4d6e-bf88-52694d5cc6c8 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning, January 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning, January 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:30.155577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:25.709602Z digest=sha256:ac3482c2f8ddd6e9fdc706e27312add6ef68817cc55386d8b275f04ed5d774d4

Observation 5bfd9ae4-d6fb-48ab-9074-776f40b33358 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Reward Model Generalization for Compute-Aware Test-Time Reasoning RLHF Workflow: From Reward Modeling to Online RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:25.802972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:25.802972Z digest=sha256:f1a43462480614279a9f0738e1fd83fe5ea9cd97dda56898ee40f9cfa8f5941f

Observation 9b4d4216-0c01-49c7-bc2e-deb80b3ec527 · outbound

This paper cites Skywork-o1 open series.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Skywork-o1 open series

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.991155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:25.912177Z digest=sha256:b6091819072479acecd7c4adc2b8b64b82d5fafc9e917b9c2664768ae5c31625

Observation 63e3dcf3-98ff-48b8-b892-fd046de0b922 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models, April 2025.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models, April 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.885755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.045623Z digest=sha256:6ecb71864e8ece5d2e12abae2e4762a7aaffa85fb987fe689eda69528010c47b

Observation b298c93a-8765-4701-b57a-9268218fcdb4 · outbound

This paper cites Self- Refine: Iterative Refinement with Self-Feedback.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self- Refine: Iterative Refinement with Self-Feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.779220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.134417Z digest=sha256:27acae37d847c3436846c6e823e63f8028ceac0d1573add26c0c24371d0a4040

Observation 79eae5de-7d27-441b-9245-a15d3ef2d0c8 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self-critiquing models for assisting human evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.240819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.240819Z digest=sha256:0e3081efa14c6a8faeb91bc7bba428b09f2f0f48db49abd260061a9ae0567281

Observation 4ab4ba77-f6a9-464a-a491-97db88620f98 · outbound

This paper cites an unresolved cited work.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:29.651910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.313976Z digest=sha256:023246d92e478939d05152d7eb4068cf93dd0f90b7f47990d1e58c56dbf67388

Observation d1580805-140a-400e-af05-3e50ae1c20c2 · outbound

This paper cites Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.538636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.414898Z digest=sha256:bdd023228a7d99602b9fc585d49d87acdb0d368c834832326dd8d061ca734feb

Observation 23457b43-4bf6-4ab9-986a-95249c5d837d · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models, October 2022.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Automatic Chain of Thought Prompting in Large Language Models, October 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.403676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.484880Z digest=sha256:c813bdb81aa0abb67d67b2163ff73d9976c7f441096706b8ba2eb1de4bdd6e54

Observation 3e53c68c-5b28-483c-8688-904cf476686a · outbound

This paper cites Le, Christopher Ré, and Azalia Mirhoseini.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, Christopher Ré, and Azalia Mirhoseini

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.235347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.586366Z digest=sha256:41865277296acee481d947fa9484d9a5b3ea5d34ef98727ac86a32e67f6aa011

Observation 80aeda13-4e57-474a-830d-9b7989e22c95 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Solving math word problems with process- and outcome-based feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.655098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.655098Z digest=sha256:4709d5f8e37d4487d3cf8c03d8ad581ef1134e70b68b0ebad6a61771543255c6

Observation 3f49e629-983f-402e-8de3-c6bd1a6974a4 · outbound

This paper cites Self-Evaluation Guided Beam Search for Reasoning, October 2023.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Self-Evaluation Guided Beam Search for Reasoning, October 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:29.066930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.752346Z digest=sha256:5fed7189bff58c497364d23f1e60c4cf6fbd28405600e0d2defc0c2057cd81c5

Observation 5989dbf9-6351-427a-96f7-a2b83417e589 · outbound

This paper cites Le, Ed H.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Le, Ed H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.888280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.821415Z digest=sha256:a42da24f10a548fd9a43a5d13ba04867b9a1a3e347ce90823f1b69ba8065ce96

Observation 0f63f736-7ba2-4871-bd45-e37bdb4fc507 · outbound

This paper cites Sparsity-aware generalization theory for deep neural networks.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Sparsity-aware generalization theory for deep neural networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.710906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.878196Z digest=sha256:c42452e7ed18010369494a7233f1f2cbca2df2a1f96886ca898731dd6773e9d1

Observation bcd5bd05-065f-458f-8156-47c1acb6cf61 · outbound

This paper cites Pac-bayes compression bounds so tight that they can explain generalization.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayes compression bounds so tight that they can explain generalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.529245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:26.941110Z digest=sha256:9c7e8a479ee6b5f187b9b2dc3bb6e730c264420e053afc65a12534c6ba875877

Observation 4310f865-1e8a-40e8-8302-6476b744c6a7 · outbound

This paper cites Efficient content-based sparse attention with routing transformers.Transactions of the Association for Computational Linguistics, 9:53–68, 2021.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Efficient content-based sparse attention with routing transformers.Transactions of the Association for Computational Linguistics, 9:53–68, 2021

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.981219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.981219Z digest=sha256:dd2eb001c7af1cd964e1902ceac92a51978d9107c8ee7795aa1d1341164fb4e2

Observation aeec097a-a6a7-4685-b0a1-af1c8d334661 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Reward Model Generalization for Compute-Aware Test-Time Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.030634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.030634Z digest=sha256:191a5d653c632fe44355e8b6834bfe325517339c574ba68b8271bbee05e0aee6

Observation ae4a9032-b346-4db8-893c-4db624af8ef8 · outbound

This paper cites Policy gradient meth- ods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Policy gradient meth- ods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.086932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.086932Z digest=sha256:3114263a81ac62b708b88026db40cd7fa1c47b3ce8e8cc1c67a93a8aa5d2618a

Observation 8ab56fe0-8d44-4325-bf39-5dde9dc87c86 · outbound

This paper cites Pac-bayesian model averaging.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayesian model averaging

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.135988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.135988Z digest=sha256:483ca59cb8f321d89303b04b6d28249a08f9ce0526b1103529790ca1bf1af5a6

Observation 908e7d2b-f6af-4cf3-ae26-283ddea24f6a · outbound

This paper cites Pac-bayesian generalisation error bounds for gaussian process classification.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Pac-bayesian generalisation error bounds for gaussian process classification

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:28.359430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:27.172308Z digest=sha256:e5d64ef93d6c76889542aa47b3565cc35d888bc8d1d7ba1477f75b2192497806

Observation bd957e38-9971-47c0-9ac3-95acd361c0e5 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Training Verifiers to Solve Math Word Problems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.222407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.222407Z digest=sha256:febab0c9418a3b1f2e663b678d6a2d84e844f7de6e1aba8fca4c7b2643278efa

Observation 88c52a6c-370f-4d0c-8894-c2d94f5e85d4 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.263006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.263006Z digest=sha256:ea96cef91b5adfe3a70eac558542a041392604d1050a9a03aab9f782a30001bb

Observation f9b003d7-32fe-4b32-ab05-8497d578fab5 · outbound

This paper cites Enhancing LLM Reasoning with Reward-guided Tree Search.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Enhancing LLM Reasoning with Reward-guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.309575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.309575Z digest=sha256:347d25fb48bca8cd58f66b60adfab31f4135901fde0fc48e404403ca9dea9f02

Observation ff489b9b-8686-4b99-929b-ebd174b217a5 · outbound

This paper cites LiteSearch: Efficacious Tree Search for LLM.

Reward Model Generalization for Compute-Aware Test-Time Reasoning LiteSearch: Efficacious Tree Search for LLM

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.388786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.388786Z digest=sha256:869d217a27149b62e1c7d2daeea73f818f8d3cab1a2faf11b65a7939b502a97d

Observation 5deda4ba-41c6-44cd-8c92-812854e08f0d · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Reward Model Generalization for Compute-Aware Test-Time Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.432302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.432302Z digest=sha256:00498f75356fccfaa1732ea6a73ac4ead450f77f23933192c59483d19bd8a8b2

Observation b0e8a27f-7ab7-4583-9ee6-bbe4a596520d · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.461851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.461851Z digest=sha256:8116275a6c262e45a6163f907f8a5cf778b186bb72c5b032ad486cf6f0d14728

Observation 4ded9f65-0ab6-491b-815f-f82eeb5b2c6c · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.488154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.488154Z digest=sha256:d08213b58228693eb80bbbe40f85596aebd04ac8ada1b880de0b39ec2a9e6c7c

Observation b23991e7-dce8-4463-9a82-445e0a635091 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.539519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.539519Z digest=sha256:df2f4c9b58222b2851a0c7d0043a42a913669f006d3b5a73b0dfe45698503c8f

Observation 534d5fd5-3165-403d-b888-4f26d9ddf483 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.571108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.571108Z digest=sha256:0e41344bacf770e92959af4fed6cae184fe2b6ac487816d32147ae314b0b9c13

Observation 55695bc1-c234-4fde-aef1-1f7a950991d9 · outbound

This paper cites BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning.

Reward Model Generalization for Compute-Aware Test-Time Reasoning BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.601318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.601318Z digest=sha256:477d84ec8bf0341f32f1200e6c1a6c84a4528bfc3e810254e99fe5e3dc2497c0

Observation da4596c2-6956-4a5a-8812-5bfe8b0fe509 · outbound

This paper cites Subtracting from 1 yields the bound Equation 6.

Reward Model Generalization for Compute-Aware Test-Time Reasoning Subtracting from 1 yields the bound Equation 6

Reference 49

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:42:27.897972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:42:27.675920Z digest=sha256:aadf4233f6b57fee519b5e857dd655cfcc86d1a6ff3d627a908c69d31e0012b9

Pith citing papers

No inbound Pith citation observations are available.