Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 60 inbound Pith citation observations for arXiv:2310.03716.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:15:15.970037Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation fd3428c7-f9a8-4efb-8e9a-a8e07503d536 · inbound
Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework A Long Way to Go: Investigating Length Correlations in RLHF
Reference 176
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bda6611a-462d-4a5e-a4f9-bfdf8590a683 · inbound
Interpreting Language Reward Models via Contrastive Explanations A Long Way to Go: Investigating Length Correlations in RLHF
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34899596-ca26-4c55-96fb-253fc2dd3352 · inbound
Self-Generated Critiques Boost Reward Modeling for Language Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022c3f76-94d8-4d24-baa9-1edd7347be15 · inbound
UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699e0a8f-d3cd-489a-b6e3-dfbdf36a16e7 · inbound
When Can Proxies Improve the Sample Complexity of Preference Learning? A Long Way to Go: Investigating Length Correlations in RLHF
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2ff56d-3b00-4d5a-b7bf-0229e9336553 · inbound
CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis A Long Way to Go: Investigating Length Correlations in RLHF
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7baf5ce8-b921-4106-9d79-cf2c0f44a7b0 · inbound
Open Problems in Machine Unlearning for AI Safety A Long Way to Go: Investigating Length Correlations in RLHF
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571b9802-8a55-474e-81c0-c2e5a4ec0f20 · inbound
Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d594349-e429-4e62-a5f7-ff968947da6e · inbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A Long Way to Go: Investigating Length Correlations in RLHF
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4974abf7-e451-4b7c-b3a2-52f85b3ce142 · inbound
Online Preference Alignment for Language Models via Count-based Exploration A Long Way to Go: Investigating Length Correlations in RLHF
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e01346b-9df2-4ca3-b2cd-0c293fc45593 · inbound
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling A Long Way to Go: Investigating Length Correlations in RLHF
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f7c52d3-b3d6-435b-9a49-90fde801f122 · inbound
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e176e0-f468-4289-ba89-d9ffca8f0654 · inbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild A Long Way to Go: Investigating Length Correlations in RLHF
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823f03aa-ab4a-4e29-9d59-69a72139f1aa · inbound
Preference learning made easy: Everything should be understood through win rate A Long Way to Go: Investigating Length Correlations in RLHF
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57e7a6a-64d3-4f8a-b014-b0291c7ecc5c · inbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bba7c405-e8ce-4887-a676-306ec8f27e85 · inbound
Reinforcement Learning from Human Feedback A Long Way to Go: Investigating Length Correlations in RLHF
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6b1f4f8d-a4ec-484e-8542-d04f3331f5f3 · inbound
MPO: Multilingual Safety Alignment via Reward Gap Optimization A Long Way to Go: Investigating Length Correlations in RLHF
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90c595a-7ad6-4f05-85fa-4eaec3bd05a3 · inbound
Multi-Domain Explainability of Preferences A Long Way to Go: Investigating Length Correlations in RLHF
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3844b5c-fc83-4018-b6a9-5f94054c4f26 · inbound
Debiasing Online Preference Learning via Preference Feature Preservation A Long Way to Go: Investigating Length Correlations in RLHF
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b442089-496d-4b5a-98ed-c8938c82a44c · inbound
Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73af5cdd-1b0c-49a1-ad4f-8d4ee91000e7 · inbound
AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control A Long Way to Go: Investigating Length Correlations in RLHF
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24e63f3c-ad87-44a4-b51f-3942077965e3 · inbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5536e894-0fde-41a3-a091-46ab96deea3a · inbound
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains A Long Way to Go: Investigating Length Correlations in RLHF
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 44618efe-8da3-4200-a188-dd2b0be30815 · inbound
Factored Causal Representation Learning for Robust Reward Modeling in RLHF A Long Way to Go: Investigating Length Correlations in RLHF
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 06669edc-9c7d-46b5-a87f-e90a72d29701 · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation A Long Way to Go: Investigating Length Correlations in RLHF
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4182d5ab-7695-4523-805c-a384433e185d · inbound
Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling A Long Way to Go: Investigating Length Correlations in RLHF
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e485b73-f283-4343-adc9-cc9062d905b9 · inbound
Same Words, Different Judgments: How Preferences Vary Across Modalities A Long Way to Go: Investigating Length Correlations in RLHF
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7d7e801c-58d9-4e73-828a-34b0499054ed · inbound
Grounded Chess Reasoning in Language Models via Master Distillation A Long Way to Go: Investigating Length Correlations in RLHF
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e706f7-5350-42c3-acef-a4f37a21c5d3 · inbound
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb6bad19-b6c0-470f-90d7-156a88a82d12 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A Long Way to Go: Investigating Length Correlations in RLHF
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f5703def-dd90-49d8-8885-be064697dd6a · inbound
Robust Reward Modeling for Large Language Models via Causal Decomposition A Long Way to Go: Investigating Length Correlations in RLHF
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 247c5220-c7e8-4d73-8e78-925f1bbfb5f3 · inbound
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction A Long Way to Go: Investigating Length Correlations in RLHF
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation df56be47-4138-4850-ae37-4b4363f25a25 · inbound
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs A Long Way to Go: Investigating Length Correlations in RLHF
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation da880bd9-6413-4078-b427-eb13109afe13 · inbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Long Way to Go: Investigating Length Correlations in RLHF
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d4a8e0f9-5531-4d4e-b214-f3b43b419a14 · inbound
Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training A Long Way to Go: Investigating Length Correlations in RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 793a6de8-4a36-483c-8c96-0eca0a88efe2 · inbound
TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination A Long Way to Go: Investigating Length Correlations in RLHF
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aacc4150-6bc4-421e-8d49-e55e0eda8a76 · inbound
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders A Long Way to Go: Investigating Length Correlations in RLHF
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba470ddb-b949-421d-9f72-ae3edd45db43 · inbound
General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 76522607-8b4f-426f-a9a5-41d61fd44997 · inbound
General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 860c78bb-120c-4aaa-ab83-a06d8f332584 · inbound
General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9c15c34c-e713-4b2c-869e-825dca8795c8 · inbound
Fine-Tuning Improves Information Conveyance in Language Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 24dd6509-284b-4608-be64-273d1da92917 · inbound
HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8624b820-36b4-4fcf-8606-1021e9b205ae · inbound
Large Language Models Hack Rewards, and Society A Long Way to Go: Investigating Length Correlations in RLHF
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fba79bae-8cef-4297-9bb8-9c4fe7ed7e51 · inbound
Human agency in initial human-AI proof formalization workflows A Long Way to Go: Investigating Length Correlations in RLHF
Reference 297
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8546286-d4ec-4628-9077-4670da119ae9 · inbound
AIP: A Graph Representation for Learning and Governing Agent Skills A Long Way to Go: Investigating Length Correlations in RLHF
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4322c478-5d0e-4084-9a60-68818d879b10 · inbound
Boosting Self-Consistency with Ranking A Long Way to Go: Investigating Length Correlations in RLHF
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 27d1cab7-fd2c-4458-88ae-476ae13a737a · inbound
PAFO: Pareto Fairness Optimization for Personalized Reward Modeling A Long Way to Go: Investigating Length Correlations in RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0733ceff-fd0a-4cca-af74-700a880fc8c7 · inbound
A Unifying Lens on Reward Uncertainty in RLHF A Long Way to Go: Investigating Length Correlations in RLHF
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8eaf8cee-75bb-42df-bc50-f16fd6e34f46 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization A Long Way to Go: Investigating Length Correlations in RLHF
Reference 263
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 567ca9a4-00b3-4f25-bd9d-3914b2a0ac2b · inbound
Are LLMs Bad at Moral Reasoning? A Long Way to Go: Investigating Length Correlations in RLHF
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 879baf7e-c75c-44cb-ac42-e2b1b5283fb9 · inbound
Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 018f614d-34df-478e-9f14-e90945783695 · inbound
Safe to Check, Unsafe to Use: Relinking at the Compression Boundary of LLM Agents A Long Way to Go: Investigating Length Correlations in RLHF
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9c2fe608-fda4-420c-983c-7e2db25faa88 · inbound
BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories A Long Way to Go: Investigating Length Correlations in RLHF
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 83f3ef1f-c4fa-4943-8d0e-a742b4e7142b · inbound
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Long Way to Go: Investigating Length Correlations in RLHF
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b4ceae58-6152-4a00-b8e1-c7389e305a3a · inbound
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Long Way to Go: Investigating Length Correlations in RLHF
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 03d1d8a0-9e02-4791-b5f4-f5df460b8a94 · inbound
Attention Limited Reward Learning A Long Way to Go: Investigating Length Correlations in RLHF
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad3a143-a4bd-4dc4-9ba6-2eac61ce336e · inbound
More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges A Long Way to Go: Investigating Length Correlations in RLHF
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3b333ee2-fdde-45d5-8f8d-ccbc3e91ba2d · inbound
Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation A Long Way to Go: Investigating Length Correlations in RLHF
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e5f632-57cd-45c6-84be-0c6db6879aa0 · inbound
RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists A Long Way to Go: Investigating Length Correlations in RLHF
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2cbe56d-c148-4c65-b989-9cd4ecd9fb20 · inbound
RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists A Long Way to Go: Investigating Length Correlations in RLHF
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.