Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:18:57.117932Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2511.02130.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:18:57.117932Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T23:14:56.115834Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T07:55:58.659031Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f5f1fd78-4430-4e0e-ac5e-84d4fb6dc945 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Ai agents as universal task solvers.arXiv preprint arXiv:2510.12066, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db28879-e705-457d-83c9-00e7e1ecf568 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Weitzman
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd9cd0bd-75a9-48d6-97a4-71af8ceeefd8 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Adaptive inference-time compute: Llms can predict if they can do better, even mid-generation, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf52c75-4cac-443f-a8b2-64846505a4cd · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Learning how hard to think: Input-adaptive allocation of lm computation, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c0d873-76fc-450a-8992-79d35b225291 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning models know when they’re right: Probing hidden states for self-verification, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53cf6160-9e1d-4a0c-9523-245d0286e0ba · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning models better express their confidence
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359544d3-9c10-4bf9-a124-4a4cd27c01d5 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Are the hidden states hiding something? testing the limits of factuality-encoding capabilities in llms, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f2f98f-f6b3-428b-b689-b550b46683a8 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning When do llms admit their mistakes? understanding the role of model belief in retraction, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a724d05-908e-4bee-8398-45c60272b5ef · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Latts: Locally adaptive test-time scaling.arXiv preprint arXiv:2509.20368, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72151895-ca52-432b-9685-02026352fb2f · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e54afe1-de7c-44d9-906e-b06c25480fb8 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa48d2f6-f530-4deb-8886-e05eb718ca7d · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Scaling LLM test-time com- pute optimally can be more effective than scaling parameters for reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c97bc6-fce8-421c-b125-72039548985e · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Adaptive test-time reasoning via reward-guided dual-phase search, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae526caf-3e4d-4de6-b8b5-eb699ef6a94b · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Large language model guided tree-of-thought, 2023
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d41821a-9f60-4e60-b9fe-846a52318582 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Demystifying chains, trees, and graphs of thoughts.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–20, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af5f92b-81d8-444e-b9a8-c22d24af9fd0 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Fractured chain-of-thought reasoning, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbb725a0-d7fd-455e-bbb9-b7d8ba2d640d · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Don’t get lost in the trees: Streamlining llm reasoning by overcoming tree search exploration pitfalls, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7edc85e4-7976-4528-b532-6a187fd755e5 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Bartoldson, Bhavya Kailkhura, Guillaume Lajoie, Glen Berseth, Nikolay Malkin, and Moksh Jain
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c9fc4d1-f4ae-4db1-8b35-db2c43c619c8 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Self-consistency improves chain of thought reasoning in language models, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee55051-6514-46e4-ae26-65709cad83f3 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b90644e-ca83-4203-95db-ef85405e3824 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Answer convergence as a signal for early stopping in reasoning, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3f106b-c4fb-46cc-bf93-d812c01ccb67 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Confidence improves self-consistency in llms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4730e179-e366-476e-9708-c5e5dc94f702 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Best-of-∞ – asymptotic performance of test-time compute, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14182385-0dc7-4179-946d-681866ef5ecf · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Reasoning at the right length: Adaptive budget forcing for efficient and accu- rate LLM inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50983a61-6450-40ed-9b1b-860b734e1e3b · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Stop when enough: Adaptive early-stopping for chain-of-thought reasoning, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 786d86cd-dad1-49bf-ac6f-aa3878570218 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Dynamic early exit in reasoning models, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90d3606-b78c-489a-81de-17a3eeb9a798 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fac33f9-86a4-40b1-9d76-1a8c282ee0ec · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Tran, Yi Tay, and Donald Metzler
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ba5b92-8e66-4071-9a31-398b90ffd545 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Learning when to plan: Efficiently allocating test-time compute for llm agents, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29133d63-4fb1-4801-b915-ed33260db601 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Can past experience accelerate llm reasoning?, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fde987-d5d5-4d4e-82bc-766c6e28def9 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning React: Synergizing reasoning and acting in language models, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 502e8052-b4a4-4805-ac2a-00ddbdb5d318 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Least-to-most prompting enables complex reasoning in large language models, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98bd29e-8f15-4604-956c-bcb3abe0727e · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models, 2023
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8de295ef-7984-4c33-bc11-e214f6843e6d · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Universal model routing for efficient llm inference, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52014968-b680-4f54-a462-cd365e95e27d · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Chen, Trevor Chow, Ishan S
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b60305-9969-42f8-904b-927d5a51531a · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca0799a4-3655-4411-b409-7ab04b4e427b · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Masrouter: Learning to route llms for multi-agent systems, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49389cfa-613c-4591-9906-3d8ea5d3bb7b · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 461006de-6502-44d5-9090-9cc057fb52ca · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning s1: Simple test-time scaling
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aeae402-b595-4ac6-9205-079dbeecfc93 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46964432-08f6-40e2-a2e4-ce7ad8318d7b · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Principles of metareasoning.Artificial Intelligence, 49(1):361– 395, 1991
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce641b3-6d8b-4131-809f-8e601a04b7b2 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning The pandora’s box problem with sequential inspections, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88aa7cf-08a9-41a3-a704-77c681a8e40b · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Deep think with confidence, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a166893a-e424-4399-b681-c8af805d54fc · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Ai agents as universal task solvers, 2025
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88938956-2678-4966-9976-67728ff3f64c · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning e1: Learning adaptive control of reasoning effort.arXiv preprint arXiv:2510.27042, 2025
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 239240b7-4603-44f3-822b-e4624ea5bcb9 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning The Gittins Index: A Design Principle for Decision-Making Under Uncertainty
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ed65eb-277e-4737-803b-9ef798223146 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Universal sequential search problems.Problems of information transmission, 9(3):265–266, 1973
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b562c7ab-0e52-477a-a83e-14e5b193da90 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Department of Energy, 1978
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a548a6c5-dc39-4541-aa40-c141511a0b3e · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Cost-aware bayesian optimization via the pandora’s box gittins index.Advances in Neural Information Processing Systems, 37:115523–115562, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18455bb8-e886-4175-b0d9-53b30b46fa39 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Qwen3 Technical Report
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b01a01f2-1f88-415d-b189-f12636591218 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70cf59c5-d901-4c73-8691-fbc6a7a6970c · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 12b — problems and solutions
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4de28b56-f962-4df9-a4b5-7642a9198805 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 12a — problems and solutions
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2219d362-364b-43fe-bfb3-b9b214c1ab7d · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 10b — problems and solutions
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4292be6a-bbac-4e26-b594-7ec140183df4 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 amc 10a — problems and solutions
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43dff967-fed9-4e4c-b231-887db9ce705a · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Solving quantitative reasoning problems with language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e53748f6-761d-4301-a84b-647029e547c8 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning Let’s verify step by step, 2023
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30257503-5cd5-49c0-9a65-d9a58e30e44c · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 aime i — problems and solutions
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e18e2b8-bc6b-42ce-b133-dc20326ce6ff · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2024 aime ii — problems and solutions
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d71ef5-bd30-4ff2-ad38-578f17acc464 · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2025 aime i — problems and solutions
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb2e03b-2cad-49f6-a32f-1fe44b5feecc · outbound
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2025 aime ii — problems and solutions
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0dec1ea-6bae-4e41-9e5c-fc8c872d320b · inbound
ExecTune: Effective Steering of Black-Box LLMs with Guide Models Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bb75c710-b8c2-4b17-a3fb-8b8755367b24 · inbound
Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.