Pith. sign in

Paper Citation Record · LEDGER

Preference learning made easy: Everything should be understood through win rate

As of 17 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2502.10505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10505 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:21.331060Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6180138-a454-4301-9e33-643e150b0bb8 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

Preference learning made easy: Everything should be understood through win rate Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:23.028818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.903864Z digest=sha256:2aa6004e268bf6df790da15dc1ef2e9dbe30160e1fbb5d39498f13e9f5b46e35

Observation 1b074722-1adc-408b-b4cb-9b3f3040dc52 · outbound

This paper cites Calibration and consistency of adversarial surrogate losses.

Preference learning made easy: Everything should be understood through win rate Calibration and consistency of adversarial surrogate losses

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:23.007001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.910320Z digest=sha256:d1655553de95a5eb59bca95d023e1a4ad569dce3e2f62555c671e34ffc9d5400

Observation 4c251466-16c3-4b63-b3ac-5d2ad985737c · outbound

This paper cites Multi-class h -consistency bounds.

Preference learning made easy: Everything should be understood through win rate Multi-class h -consistency bounds

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.983272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.916903Z digest=sha256:8b2c92eabf02f08b15126f41da2e4f2d2d58b4d6dbcb1149cec67fe49be70202

Observation 4c2cc416-e36f-4b17-94ca-93af7a38517c · outbound

This paper cites G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R.

Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.965345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.924819Z digest=sha256:f2544b8a64781e87ff0f829a2c4e74b409a449d171d0862ed9525383765f3f8c

Observation d8ab9d49-91d5-49cd-a64c-2dc628851e04 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

Preference learning made easy: Everything should be understood through win rate Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.938806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.932361Z digest=sha256:f4758fcfde6635b6549be75e665de7db8baaf2f8fc323f5a92123d917f53e36f

Observation bad98efd-8a4b-48f9-9230-63a8f1eee886 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Preference learning made easy: Everything should be understood through win rate Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.938380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.938380Z digest=sha256:bb33a35e9fc4e674797ab7eb174c70d360e9cacde2456186d42425610ae12c25

Observation 5e77688e-f12f-4247-b719-10fae9e6a0b8 · outbound

This paper cites Quantile Filtered Imitation Learning.

Preference learning made easy: Everything should be understood through win rate Quantile Filtered Imitation Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T18:32:22.012736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.946064Z digest=sha256:6142099e62e3954c5a157abcfa6dfd30178c69f500674fb8b86781ff00fcceb1

Observation 43afe68a-4ec2-412e-85dc-3230471c94a9 · outbound

This paper cites Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition.

Preference learning made easy: Everything should be understood through win rate Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.919649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.952817Z digest=sha256:dcee5df2b3888510fea07282bbab62be1a0a24f428ebed6e39290d9f5c02e2bc

Observation 9f07e96b-2159-4740-a1d1-875239dda663 · outbound

This paper cites Human Alignment of Large Language Models through Online Preference Optimisation.

Preference learning made easy: Everything should be understood through win rate Human Alignment of Large Language Models through Online Preference Optimisation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.963680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.963680Z digest=sha256:0474f056d5aa69f8a12e0ce0ca25d44dbedf07098a0ccef13fc2b7255e9848a2

Observation 8107c9da-f2c5-40f5-8b2f-330139a41e75 · outbound

This paper cites H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K.

Preference learning made easy: Everything should be understood through win rate H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.898525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.971802Z digest=sha256:ea4d7bdadcdab99c15143cd068eed29f75afd3c47c145c6fefe3368651e46a88

Observation 2902befc-3e72-43b6-acfe-88eb4cef412d · outbound

This paper cites Deep reinforcement learning from human preferences.

Preference learning made easy: Everything should be understood through win rate Deep reinforcement learning from human preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.983164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.983164Z digest=sha256:fc55feeb4f3212cd4a494f739773608f5880c252e29468b442840d0e2ffbc40a

Observation b619d067-3659-4495-ad93-2921e730fb13 · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment.

Preference learning made easy: Everything should be understood through win rate Raft: Reward ranked finetuning for generative foundation model alignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.864549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.989748Z digest=sha256:fd69030d71e0933ce57177cf7a05cc518514e62f4455f06dad6bfb92b510ee16

Observation 51dd05ae-79b1-41fc-ba02-c0308782ba0f · outbound

This paper cites Understanding dataset difficulty with v-usable information.

Preference learning made easy: Everything should be understood through win rate Understanding dataset difficulty with v-usable information

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.839675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.001625Z digest=sha256:062b367566d27eba26995717d29a5a28f214251fddd2ac2d78fc9ea23d03131e

Observation c5d841ae-b317-4623-9930-c2b9d97c89b4 · outbound

This paper cites Bonbon alignment for large language models and the sweetness of best-of-n sampling.

Preference learning made easy: Everything should be understood through win rate Bonbon alignment for large language models and the sweetness of best-of-n sampling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.814056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.008913Z digest=sha256:2b0eab4035b24c1f666ce2c4c5b0cbbf367898ff470969e262b18fbc2d727c25

Observation b64d9244-8006-498c-96d9-4831e68de4b7 · outbound

This paper cites L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N.

Preference learning made easy: Everything should be understood through win rate L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.788839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.015499Z digest=sha256:2946fe39adf3decd6a971b00636a37fcb6107df593fa83e069327b460cc5124e

Observation 5724ed26-706e-4021-a240-b524f4676b17 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Preference learning made easy: Everything should be understood through win rate Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.022389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.022389Z digest=sha256:b976ab85e27426d35d23b66799034b3784bc13db754d9a694bea58db2c9ee19c

Observation 40ae999b-8da0-42a5-b001-6d340fead544 · outbound

This paper cites $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization.

Preference learning made easy: Everything should be understood through win rate $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.028109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.028109Z digest=sha256:b2ddc7ba0ba1716844c9592aa9229b94252e5ac67f5660d49ec29e06b436cd63

Observation e3f076e6-6916-4baf-95d1-94ee9fcb0d54 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Preference learning made easy: Everything should be understood through win rate Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.033460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.033460Z digest=sha256:2b83fdb5067b2605793a959cd6a485a293df9ce1429e19c88bf536ca8c7c2f9b

Observation 52740ac5-432c-4efb-870d-81ff9b8c7ff2 · outbound

This paper cites A., Choi, Y., and Hajishirzi, H.

Preference learning made easy: Everything should be understood through win rate A., Choi, Y., and Hajishirzi, H

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.766300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.038358Z digest=sha256:4b94bcd4a66d68d46bf8775af82aba914025cf59f0f55a35b460381f58c311ee

Observation 42270018-9ef5-4abf-b241-7186c0528296 · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Preference learning made easy: Everything should be understood through win rate Towards Efficient Exact Optimization of Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.045898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.045898Z digest=sha256:c53f9874f59c5d2bbf4670a170e86efda533c9d66cbbaef8f4a18ecd03b7119d

Observation eb3b3895-c27c-4a72-9c09-74f8b2d8a384 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Preference learning made easy: Everything should be understood through win rate A Survey on Human Preference Learning for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.053077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.053077Z digest=sha256:5092b9a0617ddeb963d7c7c9997f00cd8bca4915e0d4310f3a4b81b91379ed07

Observation f6ba482e-7eed-4cf0-ba69-08f0200da95b · outbound

This paper cites A survey of reinforcement learning from human feedback.

Preference learning made easy: Everything should be understood through win rate A survey of reinforcement learning from human feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.059995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.059995Z digest=sha256:555d00cbc3de15cbc5db5dbf49ac13491ace52664f9f050e118dfe6c96319f56

Observation 15754356-b360-4f9e-948e-14bdfb38a397 · outbound

This paper cites Understanding the effects of rlhf on llm generalisation and diversity.

Preference learning made easy: Everything should be understood through win rate Understanding the effects of rlhf on llm generalisation and diversity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.729922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.066417Z digest=sha256:19dc1bc1d1982e02e49b3d865d987867b69a702116e6098479a8edeac7c26071

Observation d73ad913-60ec-4daa-947a-c504d2031a5d · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Preference learning made easy: Everything should be understood through win rate OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.072721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.072721Z digest=sha256:640ef0833640eef07e88de21b7962ff4a37c414835f5295225f4c95e3c7ff288

Observation 103738c8-633e-451f-be24-7474252aa77b · outbound

This paper cites Huggingface h4 stack exchange preference dataset.

Preference learning made easy: Everything should be understood through win rate Huggingface h4 stack exchange preference dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.710240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.080097Z digest=sha256:91d20a49dd5ad04bcc658e4b6f87f5fdfdaa3cb6005c8bd99412904755b065bc

Observation 367c276c-8757-48fc-ba82-3f127d3ae281 · outbound

This paper cites Self-Alignment with Instruction Backtranslation.

Preference learning made easy: Everything should be understood through win rate Self-Alignment with Instruction Backtranslation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.086641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.086641Z digest=sha256:e21d7333b0624573a009ad23f2df2a4f8f8f10633300c525f619b326dbc6fb8c

Observation 959bff62-81b5-428d-a8df-db9dabe0082d · outbound

This paper cites an unresolved cited work.

Preference learning made easy: Everything should be understood through win rate Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:32:22.686595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.093625Z digest=sha256:3e2b021ee94626bb827526d53be3632c04decf4ce65f11d98ac1b4a92e0832f4

Observation 14b5bd41-343a-4f2a-b5ca-7ae15d2060c1 · outbound

This paper cites Large Language Models: A Survey.

Preference learning made easy: Everything should be understood through win rate Large Language Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.100629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.100629Z digest=sha256:648522c8ff948d53c933949a969fa0a7fd1c144aa2e6bbf6f904f57ec1f389ef

Observation 74293d31-b9a8-4b0c-98c0-135fa23159a2 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.663308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.113252Z digest=sha256:681632bfac53db4b13483e89fee59325b8a9ae8437fe231dcaee93925153b849

Observation 39550b1e-bb78-4ec1-85ed-2813f2bada54 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.641721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.119438Z digest=sha256:ee2a0888e55a5279744331e6cc8b7a9fb875039e034b7d74f0d0d88dd6061198

Observation 465c40fd-57e9-44af-9f9e-309cf5c2cfec · outbound

This paper cites G., Rowland, M., Guo, Z.

Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Guo, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.621373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.125490Z digest=sha256:b9bee7a4001b8f9bd1daf6b60ea97a0ddf979552f714a18170390091ae7db657

Observation 20c613a0-c772-4fd1-86f1-899de5cc54a6 · outbound

This paper cites A., Lindsten, F., and Blei, D.

Preference learning made easy: Everything should be understood through win rate A., Lindsten, F., and Blei, D

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.601091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.131839Z digest=sha256:1b115f59b4791c6d2c9c3be75939796b2c0f90cfd552871a7863da3640e8b314

Observation 5336bf60-2085-4690-9267-135c1e51011f · outbound

This paper cites Gpt-4 technical report.

Preference learning made easy: Everything should be understood through win rate Gpt-4 technical report

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.582335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.139010Z digest=sha256:0439b4a7cfe07148e90bfe2c7f70e2e11e679930a93201c6cbbde554ab71cd66

Observation 816e2397-35a3-4c92-9c0f-9672f797b1e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Preference learning made easy: Everything should be understood through win rate Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.145340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.145340Z digest=sha256:e9fcf8fdf91cb4cb3f6f430d15472152181b8ccbf3e6353ad45afe8325384605

Observation fd5fa80e-fc39-44b0-87a1-6a881f593f98 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Preference learning made easy: Everything should be understood through win rate Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.154929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.154929Z digest=sha256:847994280acfa829825a411579c47bf1886172846b167e9b4fa7763cb38e593c

Observation 85b9a4a5-97aa-499d-8585-18d3b2556cd4 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Preference learning made easy: Everything should be understood through win rate Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.165463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.165463Z digest=sha256:7b173a620a710929b71c3bc53ec78ce6f625ad2b2ea81b2ca7449e64e0db86a8

Observation f5b4a645-6411-48d5-a1b2-a2dc0a364c3c · outbound

This paper cites From r to q^* : Your language model is secretly a q-function.

Preference learning made easy: Everything should be understood through win rate From r to q^* : Your language model is secretly a q-function

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.549800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.171588Z digest=sha256:a9125573295604442df79f8248e3cffe5fb145aba3440701ee45fa8f18d13fff

Observation 1b439a24-3a82-4014-9314-a15db84d1e18 · outbound

This paper cites D., Ermon, S., and Finn, C.

Preference learning made easy: Everything should be understood through win rate D., Ermon, S., and Finn, C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.528340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.177265Z digest=sha256:d2380080b09b6ee68e0975e59a5df4d8b2fd251f0b7b65a80c0d14d5afac1ad0

Observation 05e21823-4197-4c1b-8316-77d5cd106f45 · outbound

This paper cites Black box variational inference.

Preference learning made easy: Everything should be understood through win rate Black box variational inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.507183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.183300Z digest=sha256:0f7da4483b3b331a0b1815282d630f5336766e44edbaf4f24566abbbdd11b391

Observation 2b6830b4-152f-418b-afd5-34dffd1e055f · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

Preference learning made easy: Everything should be understood through win rate Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.189249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.189249Z digest=sha256:b1f6777788b9ea6c1b95af3afa6cec1f1e92314d7a1e3a2cdfa9f13cbfa695a4

Observation a389f205-5390-4018-9842-8d8291d01b0c · outbound

This paper cites Vanishing gradients in reinforcement finetuning of language models.

Preference learning made easy: Everything should be understood through win rate Vanishing gradients in reinforcement finetuning of language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.481360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.195177Z digest=sha256:0cef0b0ef46954d39cf6209ec858843434937f90c3215be4219140e20cc942df

Observation cacef77f-dca8-49dd-a1ec-72cd0c5e9daf · outbound

This paper cites Direct nash optimization: Teaching language models to self-improve with general preferences, 2024.

Preference learning made easy: Everything should be understood through win rate Direct nash optimization: Teaching language models to self-improve with general preferences, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.452006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.200061Z digest=sha256:9eb9dbb9f10d558c9b0278541fc555232f7f24d45c1a3c710f390eee6b3fbe67

Observation 64519d34-4c42-4d2f-a260-b89f02cfed61 · outbound

This paper cites Direct Alignment with Heterogeneous Preferences.

Preference learning made easy: Everything should be understood through win rate Direct Alignment with Heterogeneous Preferences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.206355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.206355Z digest=sha256:68bcdd3dcbfda471ecdbcf59084fdb29be1fd00ae8b050d3bfc4dd40097e58f5

Observation 823f03aa-ab4a-4e29-9d59-69a72139f1aa · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Preference learning made easy: Everything should be understood through win rate A Long Way to Go: Investigating Length Correlations in RLHF

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.215175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.215175Z digest=sha256:567032f61483d94a8b1e132e33d3e50c26803376969a9be5f9d6ad30121e4728

Observation fea3bb9f-372c-4b55-89ef-4ac8e716d64b · outbound

This paper cites Distributional preference learning: Understanding and accounting for hidden context in rlhf.

Preference learning made easy: Everything should be understood through win rate Distributional preference learning: Understanding and accounting for hidden context in rlhf

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.424473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.221926Z digest=sha256:8e97f955a132b60c34d3d7eac91881ceac27345f90018fa058904a84bd0905cb

Observation d577a9f7-535e-42c8-8c1a-7fde56c8e413 · outbound

This paper cites How to compare different loss functions and their risks.

Preference learning made easy: Everything should be understood through win rate How to compare different loss functions and their risks

Reference 46

Resolution
verified exact
doi, observed 2026-08-07T18:32:21.406280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.227381Z digest=sha256:507b31f0ce925dc77227bfabe75719bec4cc2540856779e8522d59105a25b566

Observation 3b65fc6e-bc9e-44f5-a051-1c4c21f62af3 · outbound

This paper cites Learning to summarize from human feedback.

Preference learning made easy: Everything should be understood through win rate Learning to summarize from human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.232480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.232480Z digest=sha256:1e554f066c8ac767fc66fc55bd4724da1cb8210aa0f29a4ac827369609fcfa6f

Observation a9667a6d-961a-4104-b483-ad8b2d1fbb5f · outbound

This paper cites S., and Agarwal, A.

Preference learning made easy: Everything should be understood through win rate S., and Agarwal, A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.394366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.237661Z digest=sha256:6d26883e262d49e6432ead6ad739aefbb14b888a866c41911545adb6af859db2

Observation 59644008-f9e1-4566-9f7d-bb638ceb40cd · outbound

This paper cites Preference fine-tuning of llms should leverage suboptimal, on-policy data.

Preference learning made easy: Everything should be understood through win rate Preference fine-tuning of llms should leverage suboptimal, on-policy data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.375130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.243309Z digest=sha256:fb246e01247233d9f50d66ba82a8197b770cdca2cd600a93ea29b185fec4a33c

Observation a8b1c143-4b83-49ea-bf47-4462a0fc4edb · outbound

This paper cites D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P.

Preference learning made easy: Everything should be understood through win rate D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.348188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.250115Z digest=sha256:f72d6186c371af0d93177948202932cfbf7aa8afc48ed42c0d1fd1c1258458d6

Observation 8a4e59a3-d701-4d99-98be-dd44878199ec · outbound

This paper cites Trl: Transformer reinforcement learning.

Preference learning made easy: Everything should be understood through win rate Trl: Transformer reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.322709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.259111Z digest=sha256:b46df192b84a2d93871001f1692697338603f62a8a46c4b2a7b062c2896e581b

Observation 14b72f64-ca56-407f-aee4-3b0792b5404b · outbound

This paper cites Enabling language models to implicitly learn self-improvement.

Preference learning made easy: Everything should be understood through win rate Enabling language models to implicitly learn self-improvement

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.297854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.265704Z digest=sha256:f421178e1c25f7f92cd38b88abca776153c6f727bc1ba9426175ccb36401bc15

Observation 622b1c22-0d0e-4116-9b0c-ca6b54aefe74 · outbound

This paper cites Transforming and combining rewards for aligning large language models.

Preference learning made easy: Everything should be understood through win rate Transforming and combining rewards for aligning large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.261958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.273577Z digest=sha256:e7b06b82ac0a18d65b2f7b25465fb3fe71672b34136778012da8aaec8d0c7ba5

Observation 14cf1242-2bb1-4135-80d9-aa4b0aefe821 · outbound

This paper cites Policy gradient algorithms.

Preference learning made easy: Everything should be understood through win rate Policy gradient algorithms

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.238493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.282655Z digest=sha256:5346d6139208c4af651bc559fe2b69dca4b053de50782d19ce9552416ebec977

Observation 74368c74-59ad-45f1-a388-e08c26b885b0 · outbound

This paper cites V., Murray, K., and Kim, Y.

Preference learning made easy: Everything should be understood through win rate V., Murray, K., and Kim, Y

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.201317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.289186Z digest=sha256:3de74448b0ec5ea60ad85ca04cfa7bee53a3dcbe184a5f7f3e05beec0ee648a1

Observation 97ad70a4-9085-4d57-ad7e-f3d6738836c3 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Preference learning made easy: Everything should be understood through win rate Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.300233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.300233Z digest=sha256:301e7758a8b7eaaf6ae81cd16150e66ba8c14d8a44daec0de50b2a847fd7e7e1

Observation 00f98043-ba0a-4e26-a390-a345ade3e382 · outbound

This paper cites Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J.

Preference learning made easy: Everything should be understood through win rate Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.175500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.307867Z digest=sha256:dac44ec191efb2e8615ca84c62e2a07136c0c5e6822df30d6081609082e58c91

Observation 5e8863dd-e6f4-4a5c-8ebe-e6fde31070dd · outbound

This paper cites an unresolved cited work.

Preference learning made easy: Everything should be understood through win rate Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:32:22.143704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.312919Z digest=sha256:9ab74716891a77f2870f5b3384b6f8081a407b23caa02e0c5307babba3d66dc2

Observation 862eb67f-a726-4035-9669-b97892d90489 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Preference learning made easy: Everything should be understood through win rate Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.120972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.318779Z digest=sha256:4b6a87a5f54b46439195556647fd5813c79c91a988b0c95c48a09c8a2373bfe1

Observation e9610168-ebbd-4b38-b0c3-8b7fab8f24bc · outbound

This paper cites I., and Jiao, J.

Preference learning made easy: Everything should be understood through win rate I., and Jiao, J

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.100251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.325381Z digest=sha256:0145ad9c97225ba781c7ed71dec07821f58d957627f417e4373338a37a96fb0c

Observation 4419939a-a0f5-49a7-b640-ecc4d607b2e2 · outbound

This paper cites write newline.

Preference learning made easy: Everything should be understood through win rate write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.331060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.331060Z digest=sha256:f5d67ea7181c31553c7f2a1642fb6aa7e263cdaad1d6be37a5144bbc12f9d8d5

Pith citing papers

No inbound Pith citation observations are available.