Pith. sign in

Paper Citation Record · LEDGER

A Long Way to Go: Investigating Length Correlations in RLHF

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 60 inbound Pith citation observations for arXiv:2310.03716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.03716 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 60 of 60 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:15:15.970037Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fd3428c7-f9a8-4efb-8e9a-a8e07503d536 · inbound

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework cites this paper.

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework A Long Way to Go: Investigating Length Correlations in RLHF

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:15.970037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:15:15.970037Z digest=sha256:cf18c0e111540a640a4a2ed48b589c9e1ffb6b892c681a78f12e6d69d08aa81a

Observation bda6611a-462d-4a5e-a4f9-bfdf8590a683 · inbound

Interpreting Language Reward Models via Contrastive Explanations cites this paper.

Interpreting Language Reward Models via Contrastive Explanations A Long Way to Go: Investigating Length Correlations in RLHF

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.901889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.901889Z digest=sha256:66e6e460afa9b0c2574c428360a44b992d54a8e9c1532ca9af6328d2caae1dd1

Observation 34899596-ca26-4c55-96fb-253fc2dd3352 · inbound

Self-Generated Critiques Boost Reward Modeling for Language Models cites this paper.

Self-Generated Critiques Boost Reward Modeling for Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.292523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.292523Z digest=sha256:d2adf1a622c43d8e9b3e42c9a21ca1d2d0a9a26e458e2b6d1d4ffc5fbd57788d

Observation 022c3f76-94d8-4d24-baa9-1edd7347be15 · inbound

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models cites this paper.

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T14:40:29.387244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:40:29.387244Z digest=sha256:1e5f4a0c46ba6682d65eb96ed7ef9fdcb21522a9161651970605602ce13f3c2d

Observation 699e0a8f-d3cd-489a-b6e3-dfbdf36a16e7 · inbound

When Can Proxies Improve the Sample Complexity of Preference Learning? cites this paper.

When Can Proxies Improve the Sample Complexity of Preference Learning? A Long Way to Go: Investigating Length Correlations in RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.069209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.069209Z digest=sha256:680f7f514f14aa47e395e0ebfbaa73a5d4cbfda833b41522fc4dbd0cb912c46a

Observation fb2ff56d-3b00-4d5a-b7bf-0229e9336553 · inbound

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis cites this paper.

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis A Long Way to Go: Investigating Length Correlations in RLHF

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:26:46.616014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:26:46.616014Z digest=sha256:352675fc0dbb1d44a1b5df16d66c68b65c3968761afdb95366e471753b78a29d

Observation 7baf5ce8-b921-4106-9d79-cf2c0f44a7b0 · inbound

Open Problems in Machine Unlearning for AI Safety cites this paper.

Open Problems in Machine Unlearning for AI Safety A Long Way to Go: Investigating Length Correlations in RLHF

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:10.614544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:10.614544Z digest=sha256:d4e8c93d39d21b4765f2a82a5ead6c4b6568d0a02ae5e73c82b0d03334d09865

Observation 571b9802-8a55-474e-81c0-c2e5a4ec0f20 · inbound

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models cites this paper.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.685491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.685491Z digest=sha256:2596bb314befc3b65c348986e403ce26575d10327bb15533417274514f520804

Observation 9d594349-e429-4e62-a5f7-ff968947da6e · inbound

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment cites this paper.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A Long Way to Go: Investigating Length Correlations in RLHF

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.747644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.747644Z digest=sha256:833db3fc40fb314388ac5d12d7a6604558e9e92a6eb221e5cc13f39a5b7267cd

Observation 4974abf7-e451-4b7c-b3a2-52f85b3ce142 · inbound

Online Preference Alignment for Language Models via Count-based Exploration cites this paper.

Online Preference Alignment for Language Models via Count-based Exploration A Long Way to Go: Investigating Length Correlations in RLHF

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.583705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.583705Z digest=sha256:336266502f7e86df486519d435c3a5104cbc28747c8860fa748d69d6be74a83a

Observation 5e01346b-9df2-4ca3-b2cd-0c293fc45593 · inbound

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling cites this paper.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling A Long Way to Go: Investigating Length Correlations in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.272056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.272056Z digest=sha256:f79df668903e86f04584cf1a3f8cc1e02b95c885dbe1b22eb60171028f013b0e

Observation 7f7c52d3-b3d6-435b-9a49-90fde801f122 · inbound

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models cites this paper.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.210360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.210360Z digest=sha256:5d213a1f1387e77c22173dfc8d750cd791e4249bd605d4e9caf4f8f6bc821515

Observation b5e176e0-f468-4289-ba89-d9ffca8f0654 · inbound

Copilot Arena: A Platform for Code LLM Evaluation in the Wild cites this paper.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild A Long Way to Go: Investigating Length Correlations in RLHF

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.465030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.465030Z digest=sha256:0975349367b9ced284a8c9f77cde9392c7fc108d5b358be681e57a44f8690239

Observation 823f03aa-ab4a-4e29-9d59-69a72139f1aa · inbound

Preference learning made easy: Everything should be understood through win rate cites this paper.

Preference learning made easy: Everything should be understood through win rate A Long Way to Go: Investigating Length Correlations in RLHF

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.215175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.215175Z digest=sha256:2db8926497630a70d1ae93db5b4f729bf46e9f98021658b735ae8a3fc4b40184

Observation b57e7a6a-64d3-4f8a-b014-b0291c7ecc5c · inbound

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning cites this paper.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.253420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:1eaffd32b41a90775e49f2641d6c83cae0bc9c87bdff0391b58bb1c8f46f98cf

Observation bba7c405-e8ce-4887-a676-306ec8f27e85 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback A Long Way to Go: Investigating Length Correlations in RLHF

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.260438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:f87cce4847210ac8ec631674057b31d617b1955822bda71d34e408a329bf9405

Observation 6b1f4f8d-a4ec-484e-8542-d04f3331f5f3 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization A Long Way to Go: Investigating Length Correlations in RLHF

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:31.186800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:31.186800Z digest=sha256:ee66c7ae6c430b670cfbd9aa471d50fda0fcbfeaf5071165c174db419b65e96c

Observation c90c595a-7ad6-4f05-85fa-4eaec3bd05a3 · inbound

Multi-Domain Explainability of Preferences cites this paper.

Multi-Domain Explainability of Preferences A Long Way to Go: Investigating Length Correlations in RLHF

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.431299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:05:12.431299Z digest=sha256:9f5d0e5ef0f0867343ca506b09c43def035a672c6990fa8a2531d55401302a57

Observation a3844b5c-fc83-4018-b6a9-5f94054c4f26 · inbound

Debiasing Online Preference Learning via Preference Feature Preservation cites this paper.

Debiasing Online Preference Learning via Preference Feature Preservation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.676825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.676825Z digest=sha256:ac9901774ddc4d09331cdcf340b0f19751f91dd7ae15dd8b873ed2c678fb38c5

Observation 5b442089-496d-4b5a-98ed-c8938c82a44c · inbound

Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning cites this paper.

Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:13.509910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:59:13.509910Z digest=sha256:f0f0cbba3dbe86f78d52145e578b783947294ea2ab4397dfbd7e08ade542e2f9

Observation 73af5cdd-1b0c-49a1-ad4f-8d4ee91000e7 · inbound

AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control cites this paper.

AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control A Long Way to Go: Investigating Length Correlations in RLHF

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:45.978822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:00:45.978822Z digest=sha256:367269fb818138c952a85a93241b0a4606fbdec9bfe97fb4af737beda6f6cbd5

Observation 24e63f3c-ad87-44a4-b51f-3942077965e3 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:25:45.297132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:4214ee00153b1e29053eb4429e0c9558c1696cdb049a70a00d413002c5bc9e67

Observation 5536e894-0fde-41a3-a091-46ab96deea3a · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains A Long Way to Go: Investigating Length Correlations in RLHF

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.839208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:cdf21b6b90c5082150f5328db23199ead36d2a30f0403243ef9cf46c0d96208d

Observation 44618efe-8da3-4200-a188-dd2b0be30815 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF A Long Way to Go: Investigating Length Correlations in RLHF

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.606475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:affb22af6f9f5cbeda50091e1e3564f83c97bb5204b03ce9f2fb9c43feb4bd36

Observation 06669edc-9c7d-46b5-a87f-e90a72d29701 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.349733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.349733Z digest=sha256:1591a3451c36925672cec156a8f3a15399d4bf4369b5a30b34bd410d0a0a3c5b

Observation 4182d5ab-7695-4523-805c-a384433e185d · inbound

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling cites this paper.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling A Long Way to Go: Investigating Length Correlations in RLHF

Reference 2017

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:06:24.896678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:24.896678Z digest=sha256:022bf524921460cc97b664e7094b89253772e060c23f85c554247a329b8b4920

Observation 7e485b73-f283-4343-adc9-cc9062d905b9 · inbound

Same Words, Different Judgments: How Preferences Vary Across Modalities cites this paper.

Same Words, Different Judgments: How Preferences Vary Across Modalities A Long Way to Go: Investigating Length Correlations in RLHF

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:33.025808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T19:31:51.996480Z digest=sha256:2f92ec5be5f534d730c50c249f734f02c5559927c98deb503a7f2574638bc35e

Observation 7d7e801c-58d9-4e73-828a-34b0499054ed · inbound

Grounded Chess Reasoning in Language Models via Master Distillation cites this paper.

Grounded Chess Reasoning in Language Models via Master Distillation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T21:26:43.149095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:26:43.149095Z digest=sha256:2267910fd58416c922bbb9779e2a72b73f92f7e8deeb048b18fd8dfe02574035

Observation 13e706f7-5350-42c3-acef-a4f37a21c5d3 · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:16.846395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:d0582e0970fbad0bdb18dd3eca73a475cade9d0bc453e1e5a2447281a95710ff

Observation fb6bad19-b6c0-470f-90d7-156a88a82d12 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A Long Way to Go: Investigating Length Correlations in RLHF

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.520666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:cb3dcc24036bc5dc89881c9c5593d1256f79bec4862b9041889f487cfa0a46b0

Observation f5703def-dd90-49d8-8885-be064697dd6a · inbound

Robust Reward Modeling for Large Language Models via Causal Decomposition cites this paper.

Robust Reward Modeling for Large Language Models via Causal Decomposition A Long Way to Go: Investigating Length Correlations in RLHF

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:29.614432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:50:47.902911Z digest=sha256:670100e475b20aa57345f5331688ce3f810f8f5b5eeb134758848328b37b1a61

Observation 247c5220-c7e8-4d73-8e78-925f1bbfb5f3 · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction A Long Way to Go: Investigating Length Correlations in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.556587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:2709fea827b1e9a545f93e115bac29b4f3ff6304831245e077b57cc8e0da6609

Observation df56be47-4138-4850-ae37-4b4363f25a25 · inbound

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs cites this paper.

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs A Long Way to Go: Investigating Length Correlations in RLHF

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:54:16.716341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-09T22:11:40.905814Z digest=sha256:03b69541a666ffbce134695a5d0d3f2783cf23adc300a29e54e6ed168a4a787b

Observation da880bd9-6413-4078-b427-eb13109afe13 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Long Way to Go: Investigating Length Correlations in RLHF

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.374651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:a8d1e7c20806f3fcda50c2d1ec231ba67a5d50fbfe5206f66c8d5560933b1de3

Observation d4a8e0f9-5531-4d4e-b214-f3b43b419a14 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training A Long Way to Go: Investigating Length Correlations in RLHF

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.323187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:3cc60256f6cf6c6d4f389dddae7510afcf81502ab881e95c9770caa1b4cd5a39

Observation 793a6de8-4a36-483c-8c96-0eca0a88efe2 · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination A Long Way to Go: Investigating Length Correlations in RLHF

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.351806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:49e7b273d9162e9c2018ff7b0be40349baac14223a042e69270fac23da412019

Observation aacc4150-6bc4-421e-8d49-e55e0eda8a76 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders A Long Way to Go: Investigating Length Correlations in RLHF

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.113914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:580624c9a3516f3add123a7b9e432a1b6a7758f86e80f5c1153090d58e3bd9d7

Observation ba470ddb-b949-421d-9f72-ae3edd45db43 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:48:17.700027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:31c6e2b56c91cfda71dd8f3f25869567ac9102081077421053e6a3a8ca707c18

Observation 76522607-8b4f-426f-a9a5-41d61fd44997 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.903499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:bac2bad93e3ea6e9a640fc5db062de6beff7dc859af90e5a067b9b805e83ca61

Observation 860c78bb-120c-4aaa-ab83-a06d8f332584 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:24:45.692068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:5b613b37c91f81fdd5a3ae8485a6def2349d716ca6a23a36a3be4ea3e7872bdd

Observation 9c15c34c-e713-4b2c-869e-825dca8795c8 · inbound

Fine-Tuning Improves Information Conveyance in Language Models cites this paper.

Fine-Tuning Improves Information Conveyance in Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:02:46.635308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T22:58:27.094125Z digest=sha256:67617979f453af6b0e4cc2e4c0cdc82a6f2a78fc61aae8e294d9a541b29b8dc5

Observation 24dd6509-284b-4608-be64-273d1da92917 · inbound

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models cites this paper.

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.825310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T11:32:16.166724Z digest=sha256:ada7943958ddf375187585012f13558723dc244b86c2b746fa41ed63f53e9fe2

Observation 8624b820-36b4-4fcf-8606-1021e9b205ae · inbound

Large Language Models Hack Rewards, and Society cites this paper.

Large Language Models Hack Rewards, and Society A Long Way to Go: Investigating Length Correlations in RLHF

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:26.991887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:7397acd9b9a1fff7caabf672afefdc6fa5504af84fa125cdc046fd8bf7e7445e

Observation fba79bae-8cef-4297-9bb8-9c4fe7ed7e51 · inbound

Human agency in initial human-AI proof formalization workflows cites this paper.

Human agency in initial human-AI proof formalization workflows A Long Way to Go: Investigating Length Correlations in RLHF

Reference 297

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:06:34.867868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T09:29:50.282874Z digest=sha256:2e53e9a1eb680bbb3204c228d9e20b3597c9dfe51c82109df2036a356a53cb24

Observation e8546286-d4ec-4628-9077-4670da119ae9 · inbound

AIP: A Graph Representation for Learning and Governing Agent Skills cites this paper.

AIP: A Graph Representation for Learning and Governing Agent Skills A Long Way to Go: Investigating Length Correlations in RLHF

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.772402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T05:58:37.538980Z digest=sha256:97a7efca03cdfe332ad000c742a3ae12d380c5a17eb497fcda87f7e903b8251b

Observation 4322c478-5d0e-4084-9a60-68818d879b10 · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking A Long Way to Go: Investigating Length Correlations in RLHF

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.326365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:71d391df4c6c8424cab33042d9e4092e2ca950d794587c29619164469078d7fa

Observation 27d1cab7-fd2c-4458-88ae-476ae13a737a · inbound

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling cites this paper.

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling A Long Way to Go: Investigating Length Correlations in RLHF

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:23.623629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T20:00:05.900814Z digest=sha256:1dd6dd1d44d9755836d591499eef65b270bc947ca28759762f73ed2c3462bef7

Observation 0733ceff-fd0a-4cca-af74-700a880fc8c7 · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF A Long Way to Go: Investigating Length Correlations in RLHF

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.828062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:b2660bd6144e8a8c5b3c608f208ca8d97f5b2b00a47bbbb9d9e68079409b69cd

Observation 8eaf8cee-75bb-42df-bc50-f16fd6e34f46 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization A Long Way to Go: Investigating Length Correlations in RLHF

Reference 263

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.510079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:813a3d7b8de9109667c5980e44f5a37cd6b66521273d964d67072fad2b4824be

Observation 567ca9a4-00b3-4f25-bd9d-3914b2a0ac2b · inbound

Are LLMs Bad at Moral Reasoning? cites this paper.

Are LLMs Bad at Moral Reasoning? A Long Way to Go: Investigating Length Correlations in RLHF

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.661096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T08:20:24.251540Z digest=sha256:ce96ddc92d636b711682c6e65232504444800dc439eea660241d1e74f7c6f0b0

Observation 879baf7e-c75c-44cb-ac42-e2b1b5283fb9 · inbound

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning cites this paper.

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.830653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T19:04:45.062426Z digest=sha256:fbcd1ad1ecd9760970fd5469bfb299478d399be601be7d8f7b20ecec22e3926a

Observation 018f614d-34df-478e-9f14-e90945783695 · inbound

Safe to Check, Unsafe to Use: Relinking at the Compression Boundary of LLM Agents cites this paper.

Safe to Check, Unsafe to Use: Relinking at the Compression Boundary of LLM Agents A Long Way to Go: Investigating Length Correlations in RLHF

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:29:39.123429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T13:23:54.784902Z digest=sha256:ce1c6851288404f0579c274beaeb3d68b776ebb1d2e61028d9160392a48393e1

Observation 9c2fe608-fda4-420c-983c-7e2db25faa88 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories A Long Way to Go: Investigating Length Correlations in RLHF

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.493565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:558e7fdf5ae53bd2ea4abc9f0229ee3fd632a14fad017ee43d3405c5a7fe219a

Observation 83f3ef1f-c4fa-4943-8d0e-a742b4e7142b · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.164699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:2f114b30709185ae6a582f5d3d488a356f8a48eacd2f9715210c6107985589c3

Observation b4ceae58-6152-4a00-b8e1-c7389e305a3a · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.580781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-01T06:30:27.178950Z digest=sha256:6ad3f26863e1ca8261ef94666f0e15b7e7d821288f15666533d47ec099844a4d

Observation 03d1d8a0-9e02-4791-b5f4-f5df460b8a94 · inbound

Attention Limited Reward Learning cites this paper.

Attention Limited Reward Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T16:47:52.768236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:47:52.768236Z digest=sha256:f4772455f6dda8f0b6987c388b4317b8c4b8b6bb0566b06da50ea338670905ee

Observation fad3a143-a4bd-4dc4-9ba6-2eac61ce336e · inbound

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges cites this paper.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges A Long Way to Go: Investigating Length Correlations in RLHF

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T21:25:38.638530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:5dd14587b70b3574aa8ca62f8e1afe814bf921d55c19070e1efd2381b42af134

Observation 3b333ee2-fdde-45d5-8f8d-ccbc3e91ba2d · inbound

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation cites this paper.

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T15:17:58.814427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:17:58.814427Z digest=sha256:bdf797a5c9bf4eda19e6c2913bcc40e58ae83fa08c8a806f8143f27f0b166fb0

Observation c1e5f632-57cd-45c6-84be-0c6db6879aa0 · inbound

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists cites this paper.

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists A Long Way to Go: Investigating Length Correlations in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:20.503770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:45:20.503770Z digest=sha256:c13ee1a025bc784a266861479471d4043276c23b2983a8630beb8d9ea061d296

Observation b2cbe56d-c148-4c65-b989-9cd4ecd9fb20 · inbound

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists cites this paper.

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists A Long Way to Go: Investigating Length Correlations in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T17:43:23.532685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:43:23.532685Z digest=sha256:76dda58f1f96c15f4594ad8138371d11deb64ca1af75382a57477abc34895e95