Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2406.08673.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:19.063498Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation d655d4cf-9b93-415e-84c2-e0a144d85aaf · inbound
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs HelpSteer2: Open-source dataset for training top-performing reward models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d561d878-9edc-4688-9068-72a3c8c167bc · inbound
Neon: News Entity-Interaction Extraction for Enhanced Question Answering HelpSteer2: Open-source dataset for training top-performing reward models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e5792f0-9ebf-48ac-bfaf-dcd252318112 · inbound
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd HelpSteer2: Open-source dataset for training top-performing reward models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce12127a-aba2-41c0-aedc-27fc53cc946c · inbound
Interpreting Language Reward Models via Contrastive Explanations HelpSteer2: Open-source dataset for training top-performing reward models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5277e0aa-39f4-4e3d-9cdb-0541f6c20deb · inbound
Self-Generated Critiques Boost Reward Modeling for Language Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76de97a-a0c9-4a79-9842-f6b9fa669d10 · inbound
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment HelpSteer2: Open-source dataset for training top-performing reward models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe5e228-41a1-4b00-8d29-50f13a4701e5 · inbound
Yi-Lightning Technical Report HelpSteer2: Open-source dataset for training top-performing reward models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c57ab6-079c-48fc-8754-5faf0619f4f0 · inbound
Weighted-Reward Preference Optimization for Implicit Model Fusion HelpSteer2: Open-source dataset for training top-performing reward models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c6e3985-09e4-4046-9f32-5f2eb48742c4 · inbound
ALMA: Alignment with Minimal Annotation HelpSteer2: Open-source dataset for training top-performing reward models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15725ff4-32ee-4da2-80d6-910a317613ff · inbound
Data-adaptive Safety Rules for Training Reward Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19f7cfba-beb8-46ec-a317-a83a09299070 · inbound
R.I.P.: Better Models by Survival of the Fittest Prompts HelpSteer2: Open-source dataset for training top-performing reward models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a029ccd-4f98-48ef-9ea0-a947deed1b80 · inbound
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach HelpSteer2: Open-source dataset for training top-performing reward models
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1109fac1-d0c1-485c-ba78-d9ec08f5d82d · inbound
Reinforcement Learning from Human Feedback HelpSteer2: Open-source dataset for training top-performing reward models
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b6668993-b9bf-45b2-a6e8-98e72e691130 · inbound
Reinforcement Learning from Human Feedback HelpSteer2: Open-source dataset for training top-performing reward models
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68fd61bc-e436-4825-8baf-73b7ce9743dd · inbound
Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning HelpSteer2: Open-source dataset for training top-performing reward models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caed08d2-caed-4eec-b14c-a38746b092d5 · inbound
WorldPM: Scaling Human Preference Modeling HelpSteer2: Open-source dataset for training top-performing reward models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af45c27-3863-4aec-ab9a-134255b2300e · inbound
A Systematic Analysis of Base Model Choice for Reward Modeling HelpSteer2: Open-source dataset for training top-performing reward models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f717129-d5ef-4671-9290-c40b8efb86ff · inbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge HelpSteer2: Open-source dataset for training top-performing reward models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f23228-ef50-42f8-802c-ad68b2f37369 · inbound
Discriminative Policy Optimization for Token-Level Reward Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e57198dc-bc3b-4e25-9d4e-e0450c9c8766 · inbound
TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence HelpSteer2: Open-source dataset for training top-performing reward models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad03ac79-0afa-486b-b655-aa3c4bee344f · inbound
RewardBench 2: Advancing Reward Model Evaluation HelpSteer2: Open-source dataset for training top-performing reward models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 977e5d76-02ea-4199-9eea-5075d5515872 · inbound
Pairwise Calibrated Rewards for Pluralistic Alignment HelpSteer2: Open-source dataset for training top-performing reward models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b23bd3-640a-4e84-89ff-091dd43c11d9 · inbound
PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization HelpSteer2: Open-source dataset for training top-performing reward models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb732b7-d282-465f-b157-e4328d374676 · inbound
SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b7fe163-9edf-44d3-8b08-d2fcf4f8311e · inbound
ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e077090a-7ba5-4c32-8c37-46c2b97c1ef3 · inbound
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs HelpSteer2: Open-source dataset for training top-performing reward models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a415e80b-7a3e-428c-b68d-885bbdbf7a28 · inbound
OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique HelpSteer2: Open-source dataset for training top-performing reward models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be4873b-40e3-45a8-896b-91a648f1633d · inbound
Tiny Reward Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab21d6ac-d056-44f4-9ca7-01b06844b234 · inbound
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss HelpSteer2: Open-source dataset for training top-performing reward models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abf26adc-4c7d-4de0-b06e-55ef50970bdd · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers HelpSteer2: Open-source dataset for training top-performing reward models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc0d8b94-5e21-4b50-841c-5a4a769fafd5 · inbound
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases HelpSteer2: Open-source dataset for training top-performing reward models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation db8f3726-5db8-4eb9-b9e8-983974304b8c · inbound
PolyAlign: Conditional Human-Distribution Alignment HelpSteer2: Open-source dataset for training top-performing reward models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation efbb4249-af8a-485e-903b-8c4416d37995 · inbound
PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration HelpSteer2: Open-source dataset for training top-performing reward models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 26726097-ee01-4fec-bbb3-dc053c65c315 · inbound
Step-Level Preference Learning for Generative Agents in Social Simulations HelpSteer2: Open-source dataset for training top-performing reward models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff47f6f4-94a1-4420-ae85-74801ece0ce2 · inbound
Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization HelpSteer2: Open-source dataset for training top-performing reward models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.