Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2408.11791.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:54.211281Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 29829418-edc8-4424-97e7-1a55043c0341 · inbound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f674538-c9d1-4632-97fa-bcf9647e8a61 · inbound
Reinforcement Learning from Human Feedback Critique-out-Loud Reward Models
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 73873e99-f590-43c3-8f88-36992110828b · inbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Critique-out-Loud Reward Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 053d29f9-a696-4f8b-bf2c-3e99b8a48106 · inbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Critique-out-Loud Reward Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387c1d45-c326-44e0-a065-300f432830c8 · inbound
When Slower Isn't Truer: Inverse Scaling Law of Truthfulness in Multimodal Reasoning Critique-out-Loud Reward Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50165dfc-02ec-42c5-b4e2-4589e29b9d64 · inbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Critique-out-Loud Reward Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61de2095-0d17-444c-88a6-01b3badfa06f · inbound
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns Critique-out-Loud Reward Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90199c7f-2665-4981-b314-6dc5fbad52c5 · inbound
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability Critique-out-Loud Reward Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a461104-60f6-4eef-8872-5c714efec7d4 · inbound
RewardBench 2: Advancing Reward Model Evaluation Critique-out-Loud Reward Models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2782c99-9211-4367-ad11-4a63a5f59a4f · inbound
RewardAnything: Generalizable Principle-Following Reward Models Critique-out-Loud Reward Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b801cc-3afb-4edc-95fa-edd60c593ab1 · inbound
Unlocking Recursive Thinking of LLMs: Alignment via Refinement Critique-out-Loud Reward Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7185dab-cc29-44f0-aed7-77af05663516 · inbound
CyberV: Cybernetics for Test-time Scaling in Video Understanding Critique-out-Loud Reward Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88600e08-3aad-440a-b8fe-7493dd2e2d5f · inbound
A Survey on Large Language Models for Mathematical Reasoning Critique-out-Loud Reward Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 717b82e3-abeb-4939-8f27-cbe14a2d3d99 · inbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Critique-out-Loud Reward Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54cf2e96-04ce-4fd7-abac-42db1f557aa2 · inbound
Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories Critique-out-Loud Reward Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f312f2fd-44f8-487b-a850-46a7ca56ff50 · inbound
Breaking the Myth: Can Small Models Infer Postconditions Too? Critique-out-Loud Reward Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7414c9-713f-489a-a2f9-ed4b1c21aeac · inbound
RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback Critique-out-Loud Reward Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c6ce263-51aa-46b7-9c5f-c73c385aa4e9 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Critique-out-Loud Reward Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2871b6e8-113e-4917-92d1-70c386faed27 · inbound
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Critique-out-Loud Reward Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2fd3859e-1fec-4c0e-ad7c-7d64b65f3511 · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Critique-out-Loud Reward Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b3a7623-2bc3-4be7-863b-fb344aa79ecf · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Critique-out-Loud Reward Models
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94806a1e-9d9c-4889-be33-349810620731 · inbound
Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context Critique-out-Loud Reward Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9eedcf7-ea0e-4e25-a36b-53136ad5baff · inbound
Building a Precise Video Language with Human-AI Oversight Critique-out-Loud Reward Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c06dd0f4-de11-46f0-b672-6127bd5eda1f · inbound
POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference Critique-out-Loud Reward Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3950b2ef-8a37-4881-a12a-ed5e522102e3 · inbound
Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria Critique-out-Loud Reward Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 167c9dd6-8b63-46e6-9737-97b2558b2275 · inbound
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents Critique-out-Loud Reward Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b610004c-d2ae-49c4-9b69-3783db5b7733 · inbound
Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning Critique-out-Loud Reward Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32b7c71e-b8e2-43cf-9ba2-8038e38c8035 · inbound
Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web Critique-out-Loud Reward Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44aa9132-73d4-4734-a09e-d5693b32022d · inbound
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Critique-out-Loud Reward Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 60e72776-fd66-40b5-ad67-d78891cd5a0b · inbound
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight Critique-out-Loud Reward Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ecdb5959-f451-4728-8209-00316980fc7a · inbound
Trust Region On-Policy Distillation Critique-out-Loud Reward Models
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e5fdd755-982f-4f9a-a992-a42aef3bd337 · inbound
Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics Critique-out-Loud Reward Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 519f152e-7f22-4a9f-8f18-3e4c80e4fed9 · inbound
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Critique-out-Loud Reward Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485eac96-9550-4488-868a-acade5e44020 · inbound
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Critique-out-Loud Reward Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.