Pith. sign in

Paper Citation Record · LEDGER

A Long Way to Go: Investigating Length Correlations in RLHF

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 60 inbound Pith citation observations for arXiv:2310.03716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.03716 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 60 of 60 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:15:15.970037Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fd3428c7-f9a8-4efb-8e9a-a8e07503d536 · inbound

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework cites this paper.

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework A Long Way to Go: Investigating Length Correlations in RLHF

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:15.970037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:15:15.970037Z digest=sha256:0dd2735f7209838229efd0b6fdb640c0ea0ae5ed7ac163153df5d591bd0ac769

Observation bda6611a-462d-4a5e-a4f9-bfdf8590a683 · inbound

Interpreting Language Reward Models via Contrastive Explanations cites this paper.

Interpreting Language Reward Models via Contrastive Explanations A Long Way to Go: Investigating Length Correlations in RLHF

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.901889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.901889Z digest=sha256:442d06ec3231cf65c0e814ecfcdd351a9b7e5e10479c83631fd7eeec1fe947ef

Observation 34899596-ca26-4c55-96fb-253fc2dd3352 · inbound

Self-Generated Critiques Boost Reward Modeling for Language Models cites this paper.

Self-Generated Critiques Boost Reward Modeling for Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.292523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.292523Z digest=sha256:1cf84f3f228913ab3290bbf6648844cd2bfb9f0e3a81f9a6c420384db688d45f

Observation 022c3f76-94d8-4d24-baa9-1edd7347be15 · inbound

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models cites this paper.

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T14:40:29.387244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:40:29.387244Z digest=sha256:3fe4984559b6efc8728ca9976c73ae3fdef6f4df5b32dedf380cf23658f1493b

Observation 699e0a8f-d3cd-489a-b6e3-dfbdf36a16e7 · inbound

When Can Proxies Improve the Sample Complexity of Preference Learning? cites this paper.

When Can Proxies Improve the Sample Complexity of Preference Learning? A Long Way to Go: Investigating Length Correlations in RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.069209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.069209Z digest=sha256:7a4a38f9dc5c4c0c27a9ce4dd176cacfbb237cab7b0029754f34a31d67a60478

Observation fb2ff56d-3b00-4d5a-b7bf-0229e9336553 · inbound

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis cites this paper.

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis A Long Way to Go: Investigating Length Correlations in RLHF

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:26:46.616014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:26:46.616014Z digest=sha256:ea5b2bad66946df4611d97b6d4d0a34f06e9802bd984917c806e90d6e72d0ea7

Observation 7baf5ce8-b921-4106-9d79-cf2c0f44a7b0 · inbound

Open Problems in Machine Unlearning for AI Safety cites this paper.

Open Problems in Machine Unlearning for AI Safety A Long Way to Go: Investigating Length Correlations in RLHF

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:10.614544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:10.614544Z digest=sha256:06192ee9ad4c9b382ed941d3d387c5965fc966b5d881ae07aec4bc465e97144c

Observation 571b9802-8a55-474e-81c0-c2e5a4ec0f20 · inbound

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models cites this paper.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.685491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.685491Z digest=sha256:6c3824351a5b5e83ab6098527d90e57abccc2940edd3cfd451fdc942f150c845

Observation 9d594349-e429-4e62-a5f7-ff968947da6e · inbound

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment cites this paper.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A Long Way to Go: Investigating Length Correlations in RLHF

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.747644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.747644Z digest=sha256:badf33e5116725dbdd45dca9a73a6898077b0f2fffedad47bfbc0a84e0a092a1

Observation 4974abf7-e451-4b7c-b3a2-52f85b3ce142 · inbound

Online Preference Alignment for Language Models via Count-based Exploration cites this paper.

Online Preference Alignment for Language Models via Count-based Exploration A Long Way to Go: Investigating Length Correlations in RLHF

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.583705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.583705Z digest=sha256:25b3a3bebb00e1e907cbe1a9e72227c72d7ba0b3525cd1cb3d4a66b764dd0b0a

Observation 5e01346b-9df2-4ca3-b2cd-0c293fc45593 · inbound

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling cites this paper.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling A Long Way to Go: Investigating Length Correlations in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.272056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.272056Z digest=sha256:24cb243515578cce746835947519d0156be04fb149f112d774c814f433ebe2b8

Observation 7f7c52d3-b3d6-435b-9a49-90fde801f122 · inbound

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models cites this paper.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.210360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.210360Z digest=sha256:52376600b74045f511d56f743f64f145487d2b369333c1d20e404892704ccdf6

Observation b5e176e0-f468-4289-ba89-d9ffca8f0654 · inbound

Copilot Arena: A Platform for Code LLM Evaluation in the Wild cites this paper.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild A Long Way to Go: Investigating Length Correlations in RLHF

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.465030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.465030Z digest=sha256:f4ae5e14302902c8812df344d8e59f8428a39841b31297d6595e3a2b928d08c5

Observation 823f03aa-ab4a-4e29-9d59-69a72139f1aa · inbound

Preference learning made easy: Everything should be understood through win rate cites this paper.

Preference learning made easy: Everything should be understood through win rate A Long Way to Go: Investigating Length Correlations in RLHF

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.215175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.215175Z digest=sha256:db4cf49390cf5c99e4ce97edd7f1ed6a0f9927ce8934bec22df60706d5ec4188

Observation b57e7a6a-64d3-4f8a-b014-b0291c7ecc5c · inbound

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning cites this paper.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.253420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:dbe05ee7eb22386e763227031597a44cd0f681bfba687b9513131a93dd9b2933

Observation bba7c405-e8ce-4887-a676-306ec8f27e85 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback A Long Way to Go: Investigating Length Correlations in RLHF

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.260438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:ff9d85ae69bcf9687afbde6e0ea2512af7d42326b151d05af84ef421502c6938

Observation 6b1f4f8d-a4ec-484e-8542-d04f3331f5f3 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization A Long Way to Go: Investigating Length Correlations in RLHF

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:31.186800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:31.186800Z digest=sha256:2eede4a30e47b0aa6e723e5ec7752d2dec00908af1ed16135805261f91346156

Observation c90c595a-7ad6-4f05-85fa-4eaec3bd05a3 · inbound

Multi-Domain Explainability of Preferences cites this paper.

Multi-Domain Explainability of Preferences A Long Way to Go: Investigating Length Correlations in RLHF

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.431299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:05:12.431299Z digest=sha256:a646f0a955ab941bb323299a5f270a2ddcfcae4ccc18bb53711ac0115dff5e1c

Observation a3844b5c-fc83-4018-b6a9-5f94054c4f26 · inbound

Debiasing Online Preference Learning via Preference Feature Preservation cites this paper.

Debiasing Online Preference Learning via Preference Feature Preservation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.676825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.676825Z digest=sha256:57d3b0629e3297d774247f902f16b0a7e85b88a005f50721510d0977c326cb58

Observation 5b442089-496d-4b5a-98ed-c8938c82a44c · inbound

Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning cites this paper.

Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:13.509910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:59:13.509910Z digest=sha256:710087598e1e411f829fbc8af8e47233481a0d7c36d2bf3bd1daa6e4e4850dc6

Observation 73af5cdd-1b0c-49a1-ad4f-8d4ee91000e7 · inbound

AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control cites this paper.

AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control A Long Way to Go: Investigating Length Correlations in RLHF

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:45.978822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:00:45.978822Z digest=sha256:9f15ea6755953e777689da13c1914622239d77290fb1eee44d202c63b3ad5b2e

Observation 24e63f3c-ad87-44a4-b51f-3942077965e3 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:25:45.297132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:fd2f42d341fe937dfe2d1629cb31e95394153bb86710a7637d3ac6c0281e3f3a

Observation 5536e894-0fde-41a3-a091-46ab96deea3a · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains A Long Way to Go: Investigating Length Correlations in RLHF

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.839208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:2e18fca5de19bcb3fcec705e3c91b7940de726230e3f296ef6601aea0db677ca

Observation 44618efe-8da3-4200-a188-dd2b0be30815 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF A Long Way to Go: Investigating Length Correlations in RLHF

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.606475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:7035cde90f770b70500021f9aa793b457200abf4c31f0dddaf092dd435559d65

Observation 06669edc-9c7d-46b5-a87f-e90a72d29701 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.349733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.349733Z digest=sha256:e0c76cd39ee392f23c18bfd98ceaabbe0a271ebbe7aaeadc90603451c647897c

Observation 4182d5ab-7695-4523-805c-a384433e185d · inbound

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling cites this paper.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling A Long Way to Go: Investigating Length Correlations in RLHF

Reference 2017

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:06:24.896678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:24.896678Z digest=sha256:f0e3a13f539a3de2dcbb20b692b7607570542312208f0466b47e421638793d6a

Observation 7e485b73-f283-4343-adc9-cc9062d905b9 · inbound

Same Words, Different Judgments: How Preferences Vary Across Modalities cites this paper.

Same Words, Different Judgments: How Preferences Vary Across Modalities A Long Way to Go: Investigating Length Correlations in RLHF

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:33.025808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T19:31:51.996480Z digest=sha256:9912e9222f797a3d3a4d9063730b485aa805919256c686c3c9f874c84aa24674

Observation 7d7e801c-58d9-4e73-828a-34b0499054ed · inbound

Grounded Chess Reasoning in Language Models via Master Distillation cites this paper.

Grounded Chess Reasoning in Language Models via Master Distillation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T21:26:43.149095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:26:43.149095Z digest=sha256:952ebbd7806c21092b57e27d75c0cd837011ec683c97edd74095e6c4a376193b

Observation 13e706f7-5350-42c3-acef-a4f37a21c5d3 · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:16.846395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:ce2b86e8860e72ffe1a646baafcc1e7d2d34ee4e8f84b9ab06ee2e7348bec879

Observation fb6bad19-b6c0-470f-90d7-156a88a82d12 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A Long Way to Go: Investigating Length Correlations in RLHF

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.520666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:1084946f3c81f7581000ec5ad2ba02860131f79ae513cb815c3d8061193176b6

Observation f5703def-dd90-49d8-8885-be064697dd6a · inbound

Robust Reward Modeling for Large Language Models via Causal Decomposition cites this paper.

Robust Reward Modeling for Large Language Models via Causal Decomposition A Long Way to Go: Investigating Length Correlations in RLHF

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:29.614432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:50:47.902911Z digest=sha256:4e1b34d6d71a6002d27febd868d312fd329d93f0adc94812d7174eb6343c3abb

Observation 247c5220-c7e8-4d73-8e78-925f1bbfb5f3 · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction A Long Way to Go: Investigating Length Correlations in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.556587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:4221d4778025d2c6d5f3daa969120a9c0d1e04ddd1f32a65ce0ce73e6b5c2bc1

Observation df56be47-4138-4850-ae37-4b4363f25a25 · inbound

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs cites this paper.

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs A Long Way to Go: Investigating Length Correlations in RLHF

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:54:16.716341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T22:11:40.905814Z digest=sha256:98ac80e0dc1e0fd34c1b2e4cef98fc51b595aefe394a152ba4703ef9e5df3daf

Observation da880bd9-6413-4078-b427-eb13109afe13 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Long Way to Go: Investigating Length Correlations in RLHF

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.374651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:65b97f3a73e0150408eaaa6036f72484be8cb8514c427be3757fa345d3b7975e

Observation d4a8e0f9-5531-4d4e-b214-f3b43b419a14 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training A Long Way to Go: Investigating Length Correlations in RLHF

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.323187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:192dfbc42c0d5a7823231597d52feb81831a7b13e8e8ec56eae3be834fd477b1

Observation 793a6de8-4a36-483c-8c96-0eca0a88efe2 · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination A Long Way to Go: Investigating Length Correlations in RLHF

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.351806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:edaeecd61433c501b9a67076e50e6d7aeac1b1e50131b2311f1bc5f424094a4e

Observation aacc4150-6bc4-421e-8d49-e55e0eda8a76 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders A Long Way to Go: Investigating Length Correlations in RLHF

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.113914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:9e9d7446092a1be704a81bb6114403ce8f76aa41bed7b402e2c3f9a29a1f690b

Observation ba470ddb-b949-421d-9f72-ae3edd45db43 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:48:17.700027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:317de35589c2806a6665810ba6ffa5b3245f99ef2293864692d7e90b352d9a48

Observation 76522607-8b4f-426f-a9a5-41d61fd44997 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.903499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:43d544fc1034147785e43ad4980d73b55472020d6729a3550167ec1306d73f18

Observation 860c78bb-120c-4aaa-ab83-a06d8f332584 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:24:45.692068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:339a2e0901d9c407d4f6bdd0d019758e2f24a21304d2f9922da6a254b66d1d98

Observation 9c15c34c-e713-4b2c-869e-825dca8795c8 · inbound

Fine-Tuning Improves Information Conveyance in Language Models cites this paper.

Fine-Tuning Improves Information Conveyance in Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:02:46.635308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T22:58:27.094125Z digest=sha256:9a1f40438d509f69f20c492c65cae148d492111033ee29c837d56d3114653522

Observation 24dd6509-284b-4608-be64-273d1da92917 · inbound

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models cites this paper.

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.825310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T11:32:16.166724Z digest=sha256:3ca85ff2519a0b92e9d23116b1efc3e4400e858eee3a0c43d392185aa73eaba3

Observation 8624b820-36b4-4fcf-8606-1021e9b205ae · inbound

Large Language Models Hack Rewards, and Society cites this paper.

Large Language Models Hack Rewards, and Society A Long Way to Go: Investigating Length Correlations in RLHF

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:26.991887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:a167b97ba407a920adbfafac86276c731938df5ac886d157d55886199ed56295

Observation fba79bae-8cef-4297-9bb8-9c4fe7ed7e51 · inbound

Human agency in initial human-AI proof formalization workflows cites this paper.

Human agency in initial human-AI proof formalization workflows A Long Way to Go: Investigating Length Correlations in RLHF

Reference 297

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:06:34.867868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T09:29:50.282874Z digest=sha256:e0ffde9ebc0ff2365fd3725868e9c2ecf707178f094188a2455ebc6300cab8cd

Observation e8546286-d4ec-4628-9077-4670da119ae9 · inbound

AIP: A Graph Representation for Learning and Governing Agent Skills cites this paper.

AIP: A Graph Representation for Learning and Governing Agent Skills A Long Way to Go: Investigating Length Correlations in RLHF

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.772402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T05:58:37.538980Z digest=sha256:ca484aa6dbab0df2ac2ab86d4d1cd3dd1de33d3a45c1f1cd504112bd2f477e82

Observation 4322c478-5d0e-4084-9a60-68818d879b10 · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking A Long Way to Go: Investigating Length Correlations in RLHF

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.326365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:51be79154f75a0e412e14e86881fba08aa74a6d883c23edf1761fd4781e8ec59

Observation 27d1cab7-fd2c-4458-88ae-476ae13a737a · inbound

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling cites this paper.

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling A Long Way to Go: Investigating Length Correlations in RLHF

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:23.623629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T20:00:05.900814Z digest=sha256:325b0c6e450de4510277929ba8e029e080ecf47701731eea4fd17946c93e96d4

Observation 0733ceff-fd0a-4cca-af74-700a880fc8c7 · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF A Long Way to Go: Investigating Length Correlations in RLHF

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.828062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:220f40a022d9564ae8fe0e5a73c9466aeec175b75442cc2ce399843cf0f5f9a0

Observation 8eaf8cee-75bb-42df-bc50-f16fd6e34f46 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization A Long Way to Go: Investigating Length Correlations in RLHF

Reference 263

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.510079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:b71b95e56d9499b0cc842849fa26b3c2c410bc9db1a48ab5a15d4d5aa1a61c59

Observation 567ca9a4-00b3-4f25-bd9d-3914b2a0ac2b · inbound

Are LLMs Bad at Moral Reasoning? cites this paper.

Are LLMs Bad at Moral Reasoning? A Long Way to Go: Investigating Length Correlations in RLHF

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.661096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T08:20:24.251540Z digest=sha256:ecb374c4a5d1fda68f3e1a75c445514beffdc3c5df8a9b1a66d8197d3ea8c4ab

Observation 879baf7e-c75c-44cb-ac42-e2b1b5283fb9 · inbound

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning cites this paper.

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.830653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T19:04:45.062426Z digest=sha256:196c96776329d884a3dac40925d1ac5c610e854a29a05dbe81a511d5647895e1

Observation 018f614d-34df-478e-9f14-e90945783695 · inbound

Safe to Check, Unsafe to Use: Relinking at the Compression Boundary of LLM Agents cites this paper.

Safe to Check, Unsafe to Use: Relinking at the Compression Boundary of LLM Agents A Long Way to Go: Investigating Length Correlations in RLHF

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:29:39.123429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T13:23:54.784902Z digest=sha256:f5355ba01d50c376b433a67e6de7fabe6d189383a613a45af87539c5b1d7a8a2

Observation 9c2fe608-fda4-420c-983c-7e2db25faa88 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories A Long Way to Go: Investigating Length Correlations in RLHF

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.493565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:31d1eb56a06941465fa3ac5656d932c044e9e90eea794473054e98e31596d009

Observation 83f3ef1f-c4fa-4943-8d0e-a742b4e7142b · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.164699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:b32d6339aaf8debab478cc978eb94ab4c1afe42c65d0b98d210affbaffd785a2

Observation b4ceae58-6152-4a00-b8e1-c7389e305a3a · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.580781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-01T06:30:27.178950Z digest=sha256:e9c9720de002ac67db44382d21e46d1f900f6e022bc48a65a2d3a9fdaa9cc8e4

Observation 03d1d8a0-9e02-4791-b5f4-f5df460b8a94 · inbound

Attention Limited Reward Learning cites this paper.

Attention Limited Reward Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T16:47:52.768236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:47:52.768236Z digest=sha256:eb7dde1addbd7a585a0b59fe000d007eabd8bbccbd1f2787a6822666bac36d90

Observation fad3a143-a4bd-4dc4-9ba6-2eac61ce336e · inbound

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges cites this paper.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges A Long Way to Go: Investigating Length Correlations in RLHF

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T21:25:38.638530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:3c2cbe7d396d5558d8ac5fa6ff3869bf200dd1ef158cb51483a2f59ec85af740

Observation 3b333ee2-fdde-45d5-8f8d-ccbc3e91ba2d · inbound

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation cites this paper.

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T15:17:58.814427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:17:58.814427Z digest=sha256:0ee083d009719ea43dc4f647b29a9d5a028051a17ee3af9c6774dd58cd8b40ae

Observation c1e5f632-57cd-45c6-84be-0c6db6879aa0 · inbound

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists cites this paper.

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists A Long Way to Go: Investigating Length Correlations in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:20.503770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:45:20.503770Z digest=sha256:96c682f7d636b0379e20bb574a858986a5017ec0bdc28f30821ffda475bccc7c

Observation b2cbe56d-c148-4c65-b989-9cd4ecd9fb20 · inbound

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists cites this paper.

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists A Long Way to Go: Investigating Length Correlations in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T17:43:23.532685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:43:23.532685Z digest=sha256:e7ae8846342ce0b3113f15944d201aa17461ed976d6b9d9e14f7614b2e811276