Pith. sign in

Paper Citation Record · LEDGER

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling

As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 3 inbound Pith citation observations for arXiv:2502.00814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00814 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:46:29.322067Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T09:40:58.067398Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T09:42:14.018543Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e56f5ee-7e9f-43eb-a339-7cec12a716b0 · outbound

This paper cites GPT-4 Technical Report.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.116171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.116171Z digest=sha256:204efdb35cc9303c941698e7e9b567384cfa6d0e6cd92b1ff669ad8c8f1e2d02

Observation 1e687054-2d8a-4129-afdf-829a66dce354 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.120431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.120431Z digest=sha256:26ce2ea6621a977ca3194a2652fea9344f2786603cecf00f1803277efa9fe712

Observation 4d39e875-6eee-42c7-b8b1-e0d907b9e23a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.124817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.124817Z digest=sha256:0bf4ce4fd3f685001df92480dbc72185a6ac3ff30b8f8df5bdd761d28913d17e

Observation b1851354-1fa4-40e3-b08d-319ce8a0f1b7 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.128796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.128796Z digest=sha256:28f43479ecbc386a0924320677a4662fa768a87bf34f01f8d6e2c71cc2391d26

Observation 639fa08c-a265-4ea5-8b7e-60be1ec09e59 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.132420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.132420Z digest=sha256:7c57a90faa5e6792665bcf25b153b3e7aa63ca4f0174d5473d31997cf3aaa487

Observation 895328f6-ddd7-41ce-bd66-3f7585ec7d44 · outbound

This paper cites Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, and 13 others.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, and 13 others

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:46:29.902546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.136161Z digest=sha256:b15040ed8f39951ebdcf5119262c5c87aac5cf8da2297ef14f0d2227da0f01c1

Observation 5e2afa0f-02b2-4f15-b42a-8db7b9025f3f · outbound

This paper cites Noise Contrastive Alignment of Language Models with Explicit Rewards.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Noise Contrastive Alignment of Language Models with Explicit Rewards

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.139910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.139910Z digest=sha256:3b081eb76499603ef4ed8a308b07702447f849bf9cd170b394a3e01d7c3f3922

Observation b53fb7d2-d286-4eab-9a26-13e94c5e1500 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.892194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.143706Z digest=sha256:db50c3c81f562611d58523222ef0356aa68aa27d9e0cae4cc36160d7781c4b69

Observation 90e4d899-9355-48da-9094-f0bca6d4f7ff · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.147056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.147056Z digest=sha256:4f51289dde7e21f236eb0e699b87464f980d32070ec654674e0478db3bdc55f4

Observation 01613714-7209-4ffd-9a3b-3301336890b3 · outbound

This paper cites The Llama 3 Herd of Models.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.150402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.150402Z digest=sha256:b1ba0d3524fd03472f2de80b365f248a9feeb022a1cc4ce6c5ff4506b0d61542

Observation 43b1f528-7d4a-43d9-a2b4-282d6622d2cb · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.154202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.154202Z digest=sha256:8b3faab092f5d1472e78225c7fc4c3b0f904e4e7181ea4f3c3f798b3933af5f8

Observation 8b68a731-4f52-4543-82ab-7af54c4d0614 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.874901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.158039Z digest=sha256:a171e58db3b2f3d963d746018201194cfb59036db1c37c70a8e21fefc97a1138

Observation e9650ee8-d8bd-418b-aa70-d80e0581fa81 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling KTO: Model Alignment as Prospect Theoretic Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.161466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.161466Z digest=sha256:adf22016fa9f241325bc4801da657e6b384a41f293b6700b02c7f14e7b6d4aac

Observation 32ba3a61-ea14-4c2d-a423-e7d232e4e000 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.165055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.165055Z digest=sha256:f52c6c7f564530acc6278e31c8184e73ccdca5de589e8f372eed2079e7524f41

Observation 48bd23ca-8018-44b6-bd0e-0f42d18fcf4d · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.168467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.168467Z digest=sha256:ecf7e3b5bec0b0f38630dd2c919c7359a02ddfaaa0dd580707b496d88f16a7b1

Observation f6f09260-d306-4ede-8fe5-6f7ccad15201 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.171987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.171987Z digest=sha256:8cb5f18b426cf252292fe4d704014ca3fd6f094719e0de3b9c6093bb6524cb02

Observation a113081c-5738-4ae5-b925-73739651f750 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.175463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.175463Z digest=sha256:0348fbc5b367ab2832ac722dadcad7f0ece481810ba2f0247bfab953c0684cc5

Observation b9fa1fa6-77fe-4932-b1a5-42ca979a86ca · outbound

This paper cites GPT-4o System Card.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.178757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.178757Z digest=sha256:6e6609e3e9b64b397f795c151caeddc56ce54d9a79fd4e6d4cdef0dadc652663

Observation 93666958-02ac-4f04-aad1-5e24a011bbba · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.851094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.182500Z digest=sha256:c54f3bf3f62d28d9e53c758a02d9f6392402b45c01a84edf196087c003f30b48

Observation 8189cc2d-8c2e-48a8-a87d-38025becc9b4 · outbound

This paper cites Smith, Yejin Choi, and Hannaneh Hajishirzi.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Smith, Yejin Choi, and Hannaneh Hajishirzi

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:46:29.840025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.185744Z digest=sha256:a92c82e35d31e27794c57e799659821da9534d8c911afea011e77325eaa88158

Observation 6f5e6ebf-4922-4756-ba64-7020c4f606e1 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.828400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.190322Z digest=sha256:0a174785c7ec3451723f2079b2ef3ec1ed8e1b359a3393cc19e11b00d7ca5dfe

Observation 8781ca82-6335-4fac-8dfc-b7b335b52a1c · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.193590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.193590Z digest=sha256:0a8cd70c3751d90ddc404164271f0d5201bdbee82b2534762b7051f713bec097

Observation bfd90ed0-67ab-4ecf-a737-3d6991d3750e · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling o pf, Yannic Kilcher, Dimitri von R \

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:46:29.818069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.196927Z digest=sha256:699884593fcd40a372fb9ceff52ff77c334ca1a3d5c0463ab379030987549b24

Observation c6244b72-6611-425a-9530-51d8f06b9be6 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.200234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.200234Z digest=sha256:834b78710350aaa3b55ab1b08f17c9baebfd092ff2c79871b591f2fba87d06fe

Observation d995181b-647d-4fa4-aa1f-7ba9fa2467c6 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.807852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.203615Z digest=sha256:c347f3f70c0dea393cf126b93f1fbfde3c64af11095d31896d07cb6d3e758aa8

Observation 18fe0b8c-1553-4484-875b-0d72e477049e · outbound

This paper cites The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.206702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.206702Z digest=sha256:076c77fec74798dd76c4855866bd7860f6d8ae754e52644fb90ab25e2fab08d0

Observation 9c49e95d-6763-498a-b9ae-8338a9fdac8a · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling RewardBench: Evaluating Reward Models for Language Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.210313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.210313Z digest=sha256:ab83beeadb81c0c527ad8f905f221e3ef1447df32c79d59c2c7c375224f0918e

Observation deed4d9b-04f2-4c85-aaf5-3b5df0d3df2c · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.213913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.213913Z digest=sha256:500310134b526ff1e669c9ea3cf689140518caeb38e164ea10a183d537ece5ef

Observation 93579785-6e5e-480c-98ef-d69d6e2fc840 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.797693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.217524Z digest=sha256:883a70f26cc5236ee2337e648e4bd892c9c3f818117be4573ce4b98ec41babc4

Observation f75e4305-dbf3-4c81-9146-00684ed25f74 · outbound

This paper cites LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.220861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.220861Z digest=sha256:f0ecc7f187147f938188bbc98468552fa5c9bf08f5944a911bdd7032a287cec0

Observation 86bd11ef-ef1d-460a-9c8a-e014d7947641 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.224467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.224467Z digest=sha256:efcafd53b8f317166ff4908d68c2fd988f6d8e5f3280a5cf867711d4489c9316

Observation 16dc55fe-fe52-4719-84a2-43320d48ddf1 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.228268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.228268Z digest=sha256:67d5bab2536023cadc5b6041be429bc01cb57b18770281b06c6b81671db14627

Observation 132237d9-77a9-4f51-8cea-c6efd0fd23e7 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.780132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.231882Z digest=sha256:bb4ecbf5617b34eba8346d61b649e75d2ae89f9704fbdf4ce54a339a26237e36

Observation 5bb56b2a-80fe-4d09-bf07-d307e47eea2f · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Disentangling Length from Quality in Direct Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.235461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.235461Z digest=sha256:433ddf6ec1a3faba2dcf8e5e9f27d899d79b8747f2f9fe7a160577d50da96929

Observation e1da4260-4dd8-41f0-beab-a4499c2decb1 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.242728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.242728Z digest=sha256:eb6ec7f1a60cfd5467144928fa65948afdf5c7653e7a9bd7a5b6e82eab3c3c15

Observation 74ed2454-363a-430c-b1c0-d2f248191382 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.246805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.246805Z digest=sha256:ef1feb43f9e5ae951831b5987c1d62b4e66cdef4ac1f37aa310f876add3b9875

Observation 988612c6-fb9c-4508-8a18-3d04be348795 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.762929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.250319Z digest=sha256:6e83d0cff834d155abad66eb24fb7e4e8eb85edcbec17cb01659ec3c1a6e71b0

Observation ff9b348c-7716-48ce-b735-f284b88385fd · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.253981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.253981Z digest=sha256:6cd0fe649f68b4f3f6d933a3b0d69f40040599391889ce2371665a62467f5605

Observation cd98b722-028c-4980-8079-1ceae83db87f · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling WARM: On the Benefits of Weight Averaged Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.257425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.257425Z digest=sha256:8f426bbb744c22815bd883c8309675d71ac567ba2c4a5eca62f1a85820700136

Observation 372f1231-5464-4e35-a624-9d8cdafa3132 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.260970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.260970Z digest=sha256:cc9023ff723dbd3fa856abf3f6edd27cbbaa18a2b494be427c3647a23d9d7533

Observation 28f7d9d3-3b58-4538-a665-abd9caa2bc19 · outbound

This paper cites Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.264777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.264777Z digest=sha256:6174ed1a75db8bcc0f9be453e05d4bc0ddc72e650949e94ae9097bc21d7a892b

Observation 20d7361c-5354-4e99-83d3-a0503e86928a · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.746397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.268272Z digest=sha256:ec66b60c513f53db86112472d83788d69f6755dcf946dab034ba55351968accf

Observation 5e01346b-9df2-4ca3-b2cd-0c293fc45593 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling A Long Way to Go: Investigating Length Correlations in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.272056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.272056Z digest=sha256:6e345d1c0e10a7bae360ea6e45103234875be5f603bd9da861110b2334d044e7

Observation 1566380e-2250-4489-8b79-5214692ccddc · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.735480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.275748Z digest=sha256:a5a66e96aeb38acf45c8a239a0d58256077c6b999c6d1b0ed4af482ced278c1f

Observation 51a0f741-798b-4b3a-aec1-31155760caea · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.278935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.278935Z digest=sha256:2336787338b55d454e2189dcfd35c10ff13525acd826367aab1c379d4eebb713

Observation a6c63ca5-29e6-45f7-be52-9c311174f437 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Gemini: A Family of Highly Capable Multimodal Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.282181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.282181Z digest=sha256:a7a76963ae54379e7b368fc788b691c1aa94b0246c38f5f2613d8def1086f3cf

Observation 9f4e3dc2-949e-4e33-a572-987686e0971a · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.285926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.285926Z digest=sha256:b11d4396a1750d67ce9a5aa359208b5725f181a421648a78f46879e225b678c0

Observation 59ef00cd-320d-4638-9330-b947142d06bf · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.289278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.289278Z digest=sha256:fde07180e1a1dc57ff9a28bc6e61553e89edfe83cc1745fdb396b9e789fc19dd

Observation 98cbafc3-e6d1-4391-ad86-82a413e2baba · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:46:29.712191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T17:46:29.292703Z digest=sha256:f5fce7e2843dafb2b3ea9bf12c91393498558b90e723df164d2d8481474eb688

Observation 1a0f8728-8f31-4002-8b7b-e39d4c193fba · outbound

This paper cites Qwen2 Technical Report.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Qwen2 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.295841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.295841Z digest=sha256:af1f8856a09ed9b7054135071478f2b6fc6e254ad24e0b2322f7eb92a9b28204

Observation ae705c90-d7b1-4065-86eb-f4a8027ba4af · outbound

This paper cites DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.299748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.299748Z digest=sha256:998f35ee0b9786810db681b0bcbef52447d8c0e3aaf18511c6d87e132d934220

Observation 671d4dca-2f6d-4b70-9a05-76c3579da0d1 · outbound

This paper cites Following Length Constraints in Instructions.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Following Length Constraints in Instructions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.303481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.303481Z digest=sha256:fcc9cb2d0cc96c6fc8f90c8ed056b90c16fba3ac14b8eb2461aa07c02898883b

Observation 83b96025-ddea-4fdc-8cb8-39c29ce3ccd9 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.307123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.307123Z digest=sha256:c1c8e5374882cdf3cc0f899f394541735df55fdb203648d36626b3c44f1d93c9

Observation 2c349acb-28db-4358-b755-102c3c8626a4 · outbound

This paper cites an unresolved cited work.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.310451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.310451Z digest=sha256:3087fce14e78ac64eb6a5106544654203326943da184e79eaa6d0d639960944f

Observation 37292ebc-18f4-4ea9-91dd-c8adcabab1c6 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Fine-Tuning Language Models from Human Preferences

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.313959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.313959Z digest=sha256:35c9616eebc78f6a374be7208e1d8c6658a6f10619131e24cd265427f1e4fbb7

Observation 34222efa-989c-4f61-ab43-5acd0b07059e · outbound

This paper cites online" 'onlinestring :=.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling online" 'onlinestring :=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.317766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.317766Z digest=sha256:0cae5bc054ce763cd4f70e86d174913e5aada2f89e2278018b4fda9c3b620056

Observation 5f722cb3-3e73-4e09-abb8-92656311dea5 · outbound

This paper cites write newline.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.322067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.322067Z digest=sha256:46395fdfd81189d14be570a8ef7b164d90b2d69b6cb47c5582d8dfb93b8c1189

Pith citing papers

Observation 02c42889-d3d5-4bcb-b8e2-d9ce36aaa618 · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:14.020377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:1fb6461891d53ca76c1b0cd7ca1a21cf983e257e5ff437c840e6a5d83c6dcda5

Observation 35d456f8-5298-4500-804d-5153ae5aaa8e · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.559015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:8eb2a4b62d870d582f24aff4036e91797dd7611eac14f896dd35e44a8b8ded8a

Observation 7f107fd6-152e-4728-859a-cb51dd97adeb · inbound

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs cites this paper.

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:09.497274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:11:40.905814Z digest=sha256:9b488a54ede3a050742c78926c56599999e5f3ad2573611814fe0e76353f0c24