Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.11287.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T15:49:44.016691Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:59:51.866161Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 48f26aaa-f025-40cb-8a7f-01a2f6e712ee · inbound
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps Process Reward Model with Q-Value Rankings
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71dbc219-e964-4d7b-b9f2-442416720d72 · inbound
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning Process Reward Model with Q-Value Rankings
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3df4bf-21af-4c82-87ea-c8af838614f8 · inbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Process Reward Model with Q-Value Rankings
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7443a436-cbb5-447d-b931-7f039b861a59 · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs Process Reward Model with Q-Value Rankings
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9e0e3dc3-97e1-4e1e-b2b4-f66c1636120f · inbound
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? Process Reward Model with Q-Value Rankings
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d879184b-5a0e-4540-b615-95c2dd7eec5d · inbound
FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reward Model with Q-Value Rankings
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e28e85a-a52d-48c4-81a8-bcdde58098ea · inbound
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Process Reward Model with Q-Value Rankings
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dff3c38-1b04-47eb-a071-2298d56aea64 · inbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Process Reward Model with Q-Value Rankings
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c701981-9b3a-4eff-8e56-2ae0d2d47507 · inbound
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning Process Reward Model with Q-Value Rankings
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50027bf8-af34-475d-ac06-6d1b5a69ec1e · inbound
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Process Reward Model with Q-Value Rankings
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation afac84e0-9a3c-4970-9a04-08d9b45a8c40 · inbound
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Process Reward Model with Q-Value Rankings
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 23d205d3-bf76-4552-b38b-6417f0701040 · inbound
From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding Process Reward Model with Q-Value Rankings
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 53b48197-71da-42f8-9961-2a058b7e843c · inbound
Reinforcement Learning without Ground-Truth Solutions can Improve LLMs Process Reward Model with Q-Value Rankings
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.