Pith. sign in

Paper Citation Record · LEDGER

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

As of 17 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2511.02130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.02130 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:18:57.117932Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:14:56.115834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T07:55:58.659031Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5f1fd78-4430-4e0e-ac5e-84d4fb6dc945 · outbound

This paper cites Ai agents as universal task solvers.arXiv preprint arXiv:2510.12066, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Ai agents as universal task solvers.arXiv preprint arXiv:2510.12066, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.533694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.533694Z digest=sha256:3c260d64b58a554f300d3c45b45927c1b2217a44265c094d2e731ab9d6f503e5

Observation 1db28879-e705-457d-83c9-00e7e1ecf568 · outbound

This paper cites Weitzman.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Weitzman

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.599504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.599504Z digest=sha256:8bbc7519b464f7275bf5152ce7d596d62ea65ecebf64555b45df83f22d5068a2

Observation bd9cd0bd-75a9-48d6-97a4-71af8ceeefd8 · outbound

This paper cites Adaptive inference-time compute: Llms can predict if they can do better, even mid-generation, 2024.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Adaptive inference-time compute: Llms can predict if they can do better, even mid-generation, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.699031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.699031Z digest=sha256:2bc6aa78ca0cabae6a1d7157241f5ded2851553719a37adb5135f42a47b2aafe

Observation ccf52c75-4cac-443f-a8b2-64846505a4cd · outbound

This paper cites Learning how hard to think: Input-adaptive allocation of lm computation, 2024.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Learning how hard to think: Input-adaptive allocation of lm computation, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.775735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.775735Z digest=sha256:212a4935c8eefb3ff9f555897f3c22f21b2c0e43c898734249a5ea83cc31e266

Observation 53c0d873-76fc-450a-8992-79d35b225291 · outbound

This paper cites Reasoning models know when they’re right: Probing hidden states for self-verification, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning models know when they’re right: Probing hidden states for self-verification, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.852359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.852359Z digest=sha256:8e2d42ec654c17ba916e1e9d0f31263b50bfb6f720c1f8c28ae65a1c6dc9da51

Observation 53cf6160-9e1d-4a0c-9523-245d0286e0ba · outbound

This paper cites Reasoning models better express their confidence.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning models better express their confidence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:49.927869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:49.927869Z digest=sha256:23b5f154745893dad57b97c3770d392b921a8227b04d76457f117ba7ec456cb5

Observation 359544d3-9c10-4bf9-a124-4a4cd27c01d5 · outbound

This paper cites Are the hidden states hiding something? testing the limits of factuality-encoding capabilities in llms, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Are the hidden states hiding something? testing the limits of factuality-encoding capabilities in llms, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.016662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.016662Z digest=sha256:73e8cfe82c64ef1a4b0cf0fb8b517a35ffbd449838297065824bed2d2a5efac3

Observation 34f2f98f-f6b3-428b-b689-b550b46683a8 · outbound

This paper cites When do llms admit their mistakes? understanding the role of model belief in retraction, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning When do llms admit their mistakes? understanding the role of model belief in retraction, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.155089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.155089Z digest=sha256:0bb5254860955b6ec313f01382adf4e5a0f66140c70a2262828b94c20dc64476

Observation 5a724d05-908e-4bee-8398-45c60272b5ef · outbound

This paper cites Latts: Locally adaptive test-time scaling.arXiv preprint arXiv:2509.20368, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Latts: Locally adaptive test-time scaling.arXiv preprint arXiv:2509.20368, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.231644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.231644Z digest=sha256:2cb70f14065c55f2eb5364a28dc83c043d6912efe06c47171bf19062db976fe0

Observation 72151895-ca52-432b-9685-02026352fb2f · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.369476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.369476Z digest=sha256:68ebbcf18776b76ca8913d91a1a7b3109de39cbe3a421126c2cee75aa0cd6d36

Observation 1e54afe1-de7c-44d9-906e-b06c25480fb8 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.448154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.448154Z digest=sha256:2a5431ebcc4fa697a89751f1786d2ff306b03b698a0198579327f844d3999ab0

Observation aa48d2f6-f530-4deb-8886-e05eb718ca7d · outbound

This paper cites Scaling LLM test-time com- pute optimally can be more effective than scaling parameters for reasoning.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Scaling LLM test-time com- pute optimally can be more effective than scaling parameters for reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.548821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.548821Z digest=sha256:b30fd1a3a47b02b5d364dbefe120dfbd1665326d7120ab3a3443da957556c349

Observation f8c97bc6-fce8-421c-b125-72039548985e · outbound

This paper cites Adaptive test-time reasoning via reward-guided dual-phase search, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Adaptive test-time reasoning via reward-guided dual-phase search, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.669264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.669264Z digest=sha256:9f123d9ca95df632127e13acefa322f5d37de8d28a664fb34c9d6d55dc5f8c68

Observation ae526caf-3e4d-4de6-b8b5-eb699ef6a94b · outbound

This paper cites Large language model guided tree-of-thought, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Large language model guided tree-of-thought, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.797975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.797975Z digest=sha256:9274e898c52e613e095b94c5039190b1a38793ee32c49c62c3e116f3c62bb3dd

Observation 5d41821a-9f60-4e60-b9fe-846a52318582 · outbound

This paper cites Demystifying chains, trees, and graphs of thoughts.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–20, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Demystifying chains, trees, and graphs of thoughts.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–20, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:50.950726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:50.950726Z digest=sha256:16900ee41ddff28025b27b371c1ba11b219b62c4034afe26cffe87bd276a6bc9

Observation 1af5f92b-81d8-444e-b9a8-c22d24af9fd0 · outbound

This paper cites Fractured chain-of-thought reasoning, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Fractured chain-of-thought reasoning, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.109409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.109409Z digest=sha256:f046c868eea6777fa6fa5deeaf542eb4d634e0db1911cd5b9fab25ef94cb6f97

Observation cbb725a0-d7fd-455e-bbb9-b7d8ba2d640d · outbound

This paper cites Don’t get lost in the trees: Streamlining llm reasoning by overcoming tree search exploration pitfalls, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Don’t get lost in the trees: Streamlining llm reasoning by overcoming tree search exploration pitfalls, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.200061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.200061Z digest=sha256:a67bb214c09b2777126b612bd75a57c090dc46474d45859840616e493ef7294f

Observation 7edc85e4-7976-4528-b532-6a187fd755e5 · outbound

This paper cites Bartoldson, Bhavya Kailkhura, Guillaume Lajoie, Glen Berseth, Nikolay Malkin, and Moksh Jain.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Bartoldson, Bhavya Kailkhura, Guillaume Lajoie, Glen Berseth, Nikolay Malkin, and Moksh Jain

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.332541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.332541Z digest=sha256:21202930ca3592beea87d17d6341df06fb5d7360a23fc94f70ea0f65366f0c8d

Observation 1c9fc4d1-f4ae-4db1-8b35-db2c43c619c8 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Self-consistency improves chain of thought reasoning in language models, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.452615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.452615Z digest=sha256:58e96f5a3ee1fefe4b214145a5b2fa543aa30abed8980c2dd767a4523edcdb81

Observation 5ee55051-6514-46e4-ae26-65709cad83f3 · outbound

This paper cites Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning, 2024.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.600118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.600118Z digest=sha256:7591433b3d16bc71de0a88a973c39345f8107ec591257d52326783ce7e196047

Observation 9b90644e-ca83-4203-95db-ef85405e3824 · outbound

This paper cites Answer convergence as a signal for early stopping in reasoning, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Answer convergence as a signal for early stopping in reasoning, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.746637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.746637Z digest=sha256:f6d8bf863a847143a03e914a38c7aebefc9c9922913742ce2d70797dee530c69

Observation 6e3f106b-c4fb-46cc-bf93-d812c01ccb67 · outbound

This paper cites Confidence improves self-consistency in llms.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Confidence improves self-consistency in llms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.816973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.816973Z digest=sha256:064ad86794e9af07d1148448582eb06eb00e9aab5aa5680799cc3b5d2d8b1e84

Observation 4730e179-e366-476e-9708-c5e5dc94f702 · outbound

This paper cites Best-of-∞ – asymptotic performance of test-time compute, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Best-of-∞ – asymptotic performance of test-time compute, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:51.977267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:51.977267Z digest=sha256:d70c8c3718a1ef8099dfab409f376ee0a6bce4f3f004b675d79aa9ddda7b246f

Observation 14182385-0dc7-4179-946d-681866ef5ecf · outbound

This paper cites Reasoning at the right length: Adaptive budget forcing for efficient and accu- rate LLM inference.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning at the right length: Adaptive budget forcing for efficient and accu- rate LLM inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.103823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.103823Z digest=sha256:d1ebd53c907e8b2953fa8ed7f99d8b37cec5eb160fca059ab792b17eb92f6ba1

Observation 50983a61-6450-40ed-9b1b-860b734e1e3b · outbound

This paper cites Stop when enough: Adaptive early-stopping for chain-of-thought reasoning, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Stop when enough: Adaptive early-stopping for chain-of-thought reasoning, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.246411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.246411Z digest=sha256:9d46fedbde863fae117008bf38a02744885af24f31ffae5cf8551a0663dfeae7

Observation 786d86cd-dad1-49bf-ac6f-aa3878570218 · outbound

This paper cites Dynamic early exit in reasoning models, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Dynamic early exit in reasoning models, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.376737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.376737Z digest=sha256:7d073f0c3727c130db773fa04b9bd0685b4f6e8675e81d764665e93585191e84

Observation e90d3606-b78c-489a-81de-17a3eeb9a798 · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.530409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.530409Z digest=sha256:3a70afc875a7c1a52100d22cfa1d8d3c285e38eb0e5991c2b331469511ed1e96

Observation 8fac33f9-86a4-40b1-9d76-1a8c282ee0ec · outbound

This paper cites Tran, Yi Tay, and Donald Metzler.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Tran, Yi Tay, and Donald Metzler

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.631413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.631413Z digest=sha256:a259161a71abe8291db5174ed1e57461f97e31239cc81a00e5350019461e9e27

Observation a5ba5b92-8e66-4071-9a31-398b90ffd545 · outbound

This paper cites Learning when to plan: Efficiently allocating test-time compute for llm agents, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Learning when to plan: Efficiently allocating test-time compute for llm agents, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:52.737449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:52.737449Z digest=sha256:dcb5cfbc334d5cb1fb8f2483be3749f823d4723f74487c0aa2db0c540b8892f4

Observation 29133d63-4fb1-4801-b915-ed33260db601 · outbound

This paper cites Can past experience accelerate llm reasoning?, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Can past experience accelerate llm reasoning?, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.001575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.001575Z digest=sha256:49a20770eb1146cadbf5e784fd6cb328f3f25f3ce26746c0217ab126d4bac213

Observation 42fde987-d5d5-4d4e-82bc-766c6e28def9 · outbound

This paper cites React: Synergizing reasoning and acting in language models, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning React: Synergizing reasoning and acting in language models, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.107261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.107261Z digest=sha256:17097a933a7288273452a0bd75d93e0d4c9f2afb50e2e8bce50491dc488d15c6

Observation 502e8052-b4a4-4805-ac2a-00ddbdb5d318 · outbound

This paper cites Least-to-most prompting enables complex reasoning in large language models, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Least-to-most prompting enables complex reasoning in large language models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.250671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.250671Z digest=sha256:fbcde43fc95c1fd13ae2908cdcea3d35a16ed94f1669e6d4423c705e50d4905b

Observation d98bd29e-8f15-4604-956c-bcb3abe0727e · outbound

This paper cites Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.423313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.423313Z digest=sha256:e706aa47770c76ffe8a4574a1762d6280e5aebe493b4fe4a7096c03ed04b26d8

Observation 8de295ef-7984-4c33-bc11-e214f6843e6d · outbound

This paper cites Universal model routing for efficient llm inference, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Universal model routing for efficient llm inference, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.463678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.463678Z digest=sha256:e516f89ac3894f39ab19259614124e245928fb640a5a0313dce06c7728037b31

Observation 52014968-b680-4f54-a462-cd365e95e27d · outbound

This paper cites Chen, Trevor Chow, Ishan S.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Chen, Trevor Chow, Ishan S

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.570113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.570113Z digest=sha256:867f5875194c367983a93908ada40accf2178d85555b957b3d82725e388da25c

Observation c1b60305-9969-42f8-904b-927d5a51531a · outbound

This paper cites an unresolved cited work.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.732069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.732069Z digest=sha256:2ce7ce88b5875bde73e0747a257177eebb47300aa34b8201a217986c430a6995

Observation ca0799a4-3655-4411-b409-7ab04b4e427b · outbound

This paper cites Masrouter: Learning to route llms for multi-agent systems, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Masrouter: Learning to route llms for multi-agent systems, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.845979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.845979Z digest=sha256:f4bdbf1d73d1f41ba365d636957f4e56bc48ca82a1dfc5e300f69ab7ca02d54a

Observation 49389cfa-613c-4591-9906-3d8ea5d3bb7b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:53.941103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:53.941103Z digest=sha256:6ae52d993382d52b0c24ce92c9c521e1137e099ff375049873cdadabc4320aa0

Observation 461006de-6502-44d5-9090-9cc057fb52ca · outbound

This paper cites s1: Simple test-time scaling.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning s1: Simple test-time scaling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.094174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.094174Z digest=sha256:f6eff62935edcb66785de7e6bef47663170df3ca4cf91c5f60f0951e7496d5c9

Observation 5aeae402-b595-4ac6-9205-079dbeecfc93 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.255096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.255096Z digest=sha256:4f690fc38a6749039703cdc42ada4719a5a54934fc837e71b90bc0fdcac5d9de

Observation 46964432-08f6-40e2-a2e4-ce7ad8318d7b · outbound

This paper cites Principles of metareasoning.Artificial Intelligence, 49(1):361– 395, 1991.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Principles of metareasoning.Artificial Intelligence, 49(1):361– 395, 1991

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.377086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.377086Z digest=sha256:a55b613cb1f359c9449e94aff4a51d49764595d1f98f511a681d3c12002c63a8

Observation 7ce641b3-6d8b-4131-809f-8e601a04b7b2 · outbound

This paper cites The pandora’s box problem with sequential inspections, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning The pandora’s box problem with sequential inspections, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.497296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.497296Z digest=sha256:b4b48096df3dd6f8bcf0db29d04e6759a84d68db163ae1a2d0183c43511925ce

Observation b88aa7cf-08a9-41a3-a704-77c681a8e40b · outbound

This paper cites Deep think with confidence, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Deep think with confidence, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.620933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.620933Z digest=sha256:edbf086d69c19f99efd8fe7ecfcae8e4c3f46131c43e50dc0e83baf00fa5e7c9

Observation a166893a-e424-4399-b681-c8af805d54fc · outbound

This paper cites Ai agents as universal task solvers, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Ai agents as universal task solvers, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.816578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.816578Z digest=sha256:a3522b12410d0b43468dee8c0401bf76bb0b50fafa2b00979f62d3899b9bf66f

Observation 88938956-2678-4966-9976-67728ff3f64c · outbound

This paper cites e1: Learning adaptive control of reasoning effort.arXiv preprint arXiv:2510.27042, 2025.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning e1: Learning adaptive control of reasoning effort.arXiv preprint arXiv:2510.27042, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.922794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.922794Z digest=sha256:b36012a2b3d38c66689a29928f9138095ac584ac0b029e998dfa7909603632f3

Observation 239240b7-4603-44f3-822b-e4624ea5bcb9 · outbound

This paper cites The Gittins Index: A Design Principle for Decision-Making Under Uncertainty.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning The Gittins Index: A Design Principle for Decision-Making Under Uncertainty

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.114608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.114608Z digest=sha256:b9b9d86015ec0d4be018f027dadca6fab13e21e98cd723f81420bd79f2c9db3c

Observation c6ed65eb-277e-4737-803b-9ef798223146 · outbound

This paper cites Universal sequential search problems.Problems of information transmission, 9(3):265–266, 1973.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Universal sequential search problems.Problems of information transmission, 9(3):265–266, 1973

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.255548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.255548Z digest=sha256:aadf617d1c3a51ce24d95eeb7456badccd91846e9eda5fc9b6ab39797e5a0aa3

Observation b562c7ab-0e52-477a-a83e-14e5b193da90 · outbound

This paper cites Department of Energy, 1978.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Department of Energy, 1978

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.395824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.395824Z digest=sha256:2e281e6921c66473794ddae2899457f4a27cf3a2ae63d3a882750bfbe049ce02

Observation a548a6c5-dc39-4541-aa40-c141511a0b3e · outbound

This paper cites Cost-aware bayesian optimization via the pandora’s box gittins index.Advances in Neural Information Processing Systems, 37:115523–115562, 2024.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Cost-aware bayesian optimization via the pandora’s box gittins index.Advances in Neural Information Processing Systems, 37:115523–115562, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.472999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.472999Z digest=sha256:d12636e7cafa59e203ecfff57c6e39d8bb607277b16bea36c5f0cc7c0014e008

Observation 18455bb8-e886-4175-b0d9-53b30b46fa39 · outbound

This paper cites Qwen3 Technical Report.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Qwen3 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.559048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.559048Z digest=sha256:9bfb5bc6e20693a2ed4a5bc7771e9f79f48bce9ac5b6fa065b0b986d594e07c9

Observation b01a01f2-1f88-415d-b189-f12636591218 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.692856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.692856Z digest=sha256:dbe325b1ff08563de1dcbef7523215f54b6841c216c67bacb0ac5f368c5b0841

Observation 70cf59c5-d901-4c73-8691-fbc6a7a6970c · outbound

This paper cites 2024 amc 12b — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 12b — problems and solutions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.809995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.809995Z digest=sha256:1e64bea412d72a40e19e02cefcf0ffca19bbfef8a2f917631b9cf216faaea9fb

Observation 4de28b56-f962-4df9-a4b5-7642a9198805 · outbound

This paper cites 2024 amc 12a — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 12a — problems and solutions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:55.923446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:55.923446Z digest=sha256:117b1c790ad8feee96628121902c00c1d3fc78a65917302a2f7a02beda3f7773

Observation 2219d362-364b-43fe-bfb3-b9b214c1ab7d · outbound

This paper cites 2024 amc 10b — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 10b — problems and solutions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.049118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.049118Z digest=sha256:69ce32e85629f68a30edeed9bff58d9561520b90e1a11a89dfd918b04009fbde

Observation 4292be6a-bbac-4e26-b594-7ec140183df4 · outbound

This paper cites 2024 amc 10a — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 10a — problems and solutions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.171172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.171172Z digest=sha256:95d48d8c6a2413bdc3e42dcdfda63078c6893e859bbdc6289fbe8d6b41fa0b20

Observation 43dff967-fed9-4e4c-b231-887db9ce705a · outbound

This paper cites Solving quantitative reasoning problems with language models.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Solving quantitative reasoning problems with language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.300191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.300191Z digest=sha256:8a177da73d654698ec09a95562c54d022608b65a424a94839eca912997d597a7

Observation e53748f6-761d-4301-a84b-647029e547c8 · outbound

This paper cites Let’s verify step by step, 2023.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Let’s verify step by step, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.479317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.479317Z digest=sha256:1d8fad2bd00c1d689b128a70407bc35eaf61d3f629499b0a5d4f45b60ddc5c14

Observation 30257503-5cd5-49c0-9a65-d9a58e30e44c · outbound

This paper cites 2024 aime i — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 aime i — problems and solutions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.651338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.651338Z digest=sha256:540a68dea087acc098e41ef0278ffa2733b3cc9754967f55aad052e3ae896a3f

Observation 8e18e2b8-bc6b-42ce-b133-dc20326ce6ff · outbound

This paper cites 2024 aime ii — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 aime ii — problems and solutions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.825467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.825467Z digest=sha256:3545b1170683dc3b23d137286d9b0b0d8d29eb8aea09ebc5e2e164bee92a9ec3

Observation 80d71ef5-bd30-4ff2-ad38-578f17acc464 · outbound

This paper cites 2025 aime i — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2025 aime i — problems and solutions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:56.973166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:56.973166Z digest=sha256:775c4a6a848d712f857714ef8d51d360eee4c632de43314821a293eaaf8ea15b

Observation fcb2e03b-2cad-49f6-a32f-1fe44b5feecc · outbound

This paper cites 2025 aime ii — problems and solutions.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2025 aime ii — problems and solutions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:57.117932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:57.117932Z digest=sha256:714851e1632884cee1b1a8f1aa0aa932dbbd084a513fc096fd89b014749bdb8c

Pith citing papers

Observation f0dec1ea-6bae-4e41-9e5c-fc8c872d320b · inbound

ExecTune: Effective Steering of Black-Box LLMs with Guide Models cites this paper.

ExecTune: Effective Steering of Black-Box LLMs with Guide Models Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-27T01:19:51.926662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T16:55:34.091812Z digest=sha256:07da50de8082709bef2c522414afcc692ee77febd5b36379d3a2bbe50fd9436c

Observation bb75c710-b8c2-4b17-a3fb-8b8755367b24 · inbound

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching cites this paper.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:56.115834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:56.115834Z digest=sha256:d5f52717aa23534a45a4ed0f808f35b4adff17af897c41a745eb1c7c327263c7