Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2410.09302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:20:13.735791Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T13:14:10.972580Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 929b64d2-b666-48b7-907c-3b72e02f23cc · inbound
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b41e75c3-aeea-4c77-91fd-e1d67ff4f3eb · inbound
Reinforcement Learning in hyperbolic space for multi-step reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dd6666f-6906-4848-a8cb-0064b34dee83 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a65996f4-6add-4eef-9d4f-6b316519e0b4 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc0a83a1-0b42-475e-a529-e1dede6ea1fc · inbound
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.