Pith. sign in

Paper Citation Record · LEDGER

Preference learning made easy: Everything should be understood through win rate

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2502.10505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10505 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:21.331060Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6180138-a454-4301-9e33-643e150b0bb8 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

Preference learning made easy: Everything should be understood through win rate Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:23.028818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.903864Z digest=sha256:53a6e1d5a2846e96a10801df2f3e71b27737213f299aad0dd3340381dd578814

Observation 1b074722-1adc-408b-b4cb-9b3f3040dc52 · outbound

This paper cites Calibration and consistency of adversarial surrogate losses.

Preference learning made easy: Everything should be understood through win rate Calibration and consistency of adversarial surrogate losses

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:23.007001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.910320Z digest=sha256:dc86e09c2016b60124e4ac7bc94e67ece56d7cd3da4972ba3e5964c604c7b47b

Observation 4c251466-16c3-4b63-b3ac-5d2ad985737c · outbound

This paper cites Multi-class h -consistency bounds.

Preference learning made easy: Everything should be understood through win rate Multi-class h -consistency bounds

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.983272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.916903Z digest=sha256:f34a57c48b3829d75c89c8db40ceaf36e485a1da84dc8e8756a27c37e808cb8e

Observation 4c2cc416-e36f-4b17-94ca-93af7a38517c · outbound

This paper cites G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R.

Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.965345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.924819Z digest=sha256:c619faeb435e081e9824dabcfcc04c26296b9059d1f2d649fcccd5680dcfa63a

Observation d8ab9d49-91d5-49cd-a64c-2dc628851e04 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

Preference learning made easy: Everything should be understood through win rate Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.938806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.932361Z digest=sha256:5bcc2a17e823f3e9da1be152a3eed2b356b574953ae73986e7ba661d994925e2

Observation bad98efd-8a4b-48f9-9230-63a8f1eee886 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Preference learning made easy: Everything should be understood through win rate Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.938380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.938380Z digest=sha256:58d538ef4dce5d3f3848322914d3216d44c5c4e9c2919b10604034e6eea3419f

Observation 5e77688e-f12f-4247-b719-10fae9e6a0b8 · outbound

This paper cites Quantile Filtered Imitation Learning.

Preference learning made easy: Everything should be understood through win rate Quantile Filtered Imitation Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T18:32:22.012736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.946064Z digest=sha256:ae48c5e885a64ae9ae9cfb8cbc72b83572125ea370b6e1b67b41d9dd68d2079f

Observation 43afe68a-4ec2-412e-85dc-3230471c94a9 · outbound

This paper cites Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition.

Preference learning made easy: Everything should be understood through win rate Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.919649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.952817Z digest=sha256:f8a3d0e07c672bf26193df4f9dd92a1b54779583595959dd0812e28ea95f7f73

Observation 9f07e96b-2159-4740-a1d1-875239dda663 · outbound

This paper cites Human Alignment of Large Language Models through Online Preference Optimisation.

Preference learning made easy: Everything should be understood through win rate Human Alignment of Large Language Models through Online Preference Optimisation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.963680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.963680Z digest=sha256:ad7ef22dca8128dc3f2b2994a59c2054038a32aa4932cf1a4d54ee0046e78555

Observation 8107c9da-f2c5-40f5-8b2f-330139a41e75 · outbound

This paper cites H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K.

Preference learning made easy: Everything should be understood through win rate H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.898525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.971802Z digest=sha256:6f88b46d500551c8cf919ff913caf53f08d8131a4e84632e03555b30809597af

Observation 2902befc-3e72-43b6-acfe-88eb4cef412d · outbound

This paper cites Deep reinforcement learning from human preferences.

Preference learning made easy: Everything should be understood through win rate Deep reinforcement learning from human preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.983164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.983164Z digest=sha256:cbcb36480596747e0c78f8b1d94dd4c58aac9362741d2e7a5715b89ad07034d1

Observation b619d067-3659-4495-ad93-2921e730fb13 · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment.

Preference learning made easy: Everything should be understood through win rate Raft: Reward ranked finetuning for generative foundation model alignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.864549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.989748Z digest=sha256:91afd6bb7625eea025c8b985f7179fd9fbcdceebc1a8f98a68048ca45ff1aaa7

Observation 51dd05ae-79b1-41fc-ba02-c0308782ba0f · outbound

This paper cites Understanding dataset difficulty with v-usable information.

Preference learning made easy: Everything should be understood through win rate Understanding dataset difficulty with v-usable information

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.839675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.001625Z digest=sha256:ccd4bf0c4b69f080a5d0df092a4674a745359578ac9f3c9124168f5c1ffe71d0

Observation c5d841ae-b317-4623-9930-c2b9d97c89b4 · outbound

This paper cites Bonbon alignment for large language models and the sweetness of best-of-n sampling.

Preference learning made easy: Everything should be understood through win rate Bonbon alignment for large language models and the sweetness of best-of-n sampling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.814056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.008913Z digest=sha256:77e2c2c8bbf279f37e3681754836150af47320c871d44855cdf4b65959596935

Observation b64d9244-8006-498c-96d9-4831e68de4b7 · outbound

This paper cites L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N.

Preference learning made easy: Everything should be understood through win rate L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.788839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.015499Z digest=sha256:5bf8e60beaac2cbc5babf21595e673876695db25a6f295d2e40fd25f835a35c0

Observation 5724ed26-706e-4021-a240-b524f4676b17 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Preference learning made easy: Everything should be understood through win rate Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.022389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.022389Z digest=sha256:f950d45ca21b76f2f398230e1951caa8e6ef124f0e6de254a1f9e15540357a6c

Observation 40ae999b-8da0-42a5-b001-6d340fead544 · outbound

This paper cites $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization.

Preference learning made easy: Everything should be understood through win rate $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.028109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.028109Z digest=sha256:0843bd32d5779aa08ac93458167533aa4f18a58b61d9892487c180671ef4ab5a

Observation e3f076e6-6916-4baf-95d1-94ee9fcb0d54 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Preference learning made easy: Everything should be understood through win rate Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.033460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.033460Z digest=sha256:72a467042999069da9c1f138352e359f6e6062f50d5fb42c85730891bc551956

Observation 52740ac5-432c-4efb-870d-81ff9b8c7ff2 · outbound

This paper cites A., Choi, Y., and Hajishirzi, H.

Preference learning made easy: Everything should be understood through win rate A., Choi, Y., and Hajishirzi, H

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.766300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.038358Z digest=sha256:988d1befe140f969ed74db700a1e40ba7dda1608def93d4da7a8f5a341e111cd

Observation 42270018-9ef5-4abf-b241-7186c0528296 · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Preference learning made easy: Everything should be understood through win rate Towards Efficient Exact Optimization of Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.045898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.045898Z digest=sha256:7e4c4269e49a78b847c43f7726de69165923d6af19a4f32085a1c0a9b2d94a73

Observation eb3b3895-c27c-4a72-9c09-74f8b2d8a384 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Preference learning made easy: Everything should be understood through win rate A Survey on Human Preference Learning for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.053077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.053077Z digest=sha256:194d1b01ecc63839aed86b6d188d1e6708427301b5928d235c2456ad6c3cab64

Observation f6ba482e-7eed-4cf0-ba69-08f0200da95b · outbound

This paper cites A survey of reinforcement learning from human feedback.

Preference learning made easy: Everything should be understood through win rate A survey of reinforcement learning from human feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.059995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.059995Z digest=sha256:9f9e87fff033eacf956a597e8ebebceb0963d884928fba0c7c2963cdfd795624

Observation 15754356-b360-4f9e-948e-14bdfb38a397 · outbound

This paper cites Understanding the effects of rlhf on llm generalisation and diversity.

Preference learning made easy: Everything should be understood through win rate Understanding the effects of rlhf on llm generalisation and diversity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.729922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.066417Z digest=sha256:7c02dfd724729899dce1d7a224a29e066d37b84299f006f79efc1edb313a18bc

Observation d73ad913-60ec-4daa-947a-c504d2031a5d · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Preference learning made easy: Everything should be understood through win rate OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.072721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.072721Z digest=sha256:b24de483c688d88bd14c16d859d93706e2b5b4b01523063259a0f303740fe7c4

Observation 103738c8-633e-451f-be24-7474252aa77b · outbound

This paper cites Huggingface h4 stack exchange preference dataset.

Preference learning made easy: Everything should be understood through win rate Huggingface h4 stack exchange preference dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.710240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.080097Z digest=sha256:40f40c0980443b8bb394cebc58d878519fa5f3757cb3327e583dd1e46345c3be

Observation 367c276c-8757-48fc-ba82-3f127d3ae281 · outbound

This paper cites Self-Alignment with Instruction Backtranslation.

Preference learning made easy: Everything should be understood through win rate Self-Alignment with Instruction Backtranslation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.086641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.086641Z digest=sha256:8b4a060adff06da882ee1901e3b4d3c06e7428a55a16e11c87db5e62703a98ad

Observation 959bff62-81b5-428d-a8df-db9dabe0082d · outbound

This paper cites an unresolved cited work.

Preference learning made easy: Everything should be understood through win rate Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:32:22.686595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.093625Z digest=sha256:e4ba9866db6d40a229f30a6351e0ab69510d078144ba7e108077c0e8f31e0115

Observation 14b5bd41-343a-4f2a-b5ca-7ae15d2060c1 · outbound

This paper cites Large Language Models: A Survey.

Preference learning made easy: Everything should be understood through win rate Large Language Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.100629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.100629Z digest=sha256:1ccafd43cdbc06a3cade68b0a1b5579842a3fdce2d0971dfcb606b12505d607c

Observation 74293d31-b9a8-4b0c-98c0-135fa23159a2 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.663308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.113252Z digest=sha256:a922700bd5f0338edf5cca47895e0fb10b1f28c99fce4ace532af9bc64bd48b0

Observation 39550b1e-bb78-4ec1-85ed-2813f2bada54 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.641721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.119438Z digest=sha256:8225bf45fdfddd4ad2e7c252e0a25e8bddea0ed02929e948126fdbd66a1b47b5

Observation 465c40fd-57e9-44af-9f9e-309cf5c2cfec · outbound

This paper cites G., Rowland, M., Guo, Z.

Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Guo, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.621373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.125490Z digest=sha256:42bcc99bf8da213344a318f12d40c81c261268d37532dcd22f4a3fd65e00ee95

Observation 20c613a0-c772-4fd1-86f1-899de5cc54a6 · outbound

This paper cites A., Lindsten, F., and Blei, D.

Preference learning made easy: Everything should be understood through win rate A., Lindsten, F., and Blei, D

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.601091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.131839Z digest=sha256:38af59ed3096f165eb186a0c8099e8233175f45496142cae501c9bf08db5b96c

Observation 5336bf60-2085-4690-9267-135c1e51011f · outbound

This paper cites Gpt-4 technical report.

Preference learning made easy: Everything should be understood through win rate Gpt-4 technical report

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.582335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.139010Z digest=sha256:1bb63cfd4ef68ad7533e48940ee93c29ab5a44300e31dbb930db1f4b1704628d

Observation 816e2397-35a3-4c92-9c0f-9672f797b1e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Preference learning made easy: Everything should be understood through win rate Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.145340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.145340Z digest=sha256:e01853338dadd9992fe82c7146d0b89ec0ecc063bbf1ab608652118811b6edc7

Observation fd5fa80e-fc39-44b0-87a1-6a881f593f98 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Preference learning made easy: Everything should be understood through win rate Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.154929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.154929Z digest=sha256:37111e0e22be39b811211e50576c0e6aa396e9fdd8612e02bdbb363020f07b8d

Observation 85b9a4a5-97aa-499d-8585-18d3b2556cd4 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Preference learning made easy: Everything should be understood through win rate Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.165463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.165463Z digest=sha256:16817298fe416f56c4e4b90082cc45c46e239f38034a9b561db0d4d0c62e2fed

Observation f5b4a645-6411-48d5-a1b2-a2dc0a364c3c · outbound

This paper cites From r to q^* : Your language model is secretly a q-function.

Preference learning made easy: Everything should be understood through win rate From r to q^* : Your language model is secretly a q-function

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.549800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.171588Z digest=sha256:2024e789037d8cdb83263eed05cc5f4482a9ebde649a46e74375fd30df6057bb

Observation 1b439a24-3a82-4014-9314-a15db84d1e18 · outbound

This paper cites D., Ermon, S., and Finn, C.

Preference learning made easy: Everything should be understood through win rate D., Ermon, S., and Finn, C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.528340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.177265Z digest=sha256:1bdffdb9160ff2b7b111117f9928e605e91b3fd654e3142272c339e8c817b5af

Observation 05e21823-4197-4c1b-8316-77d5cd106f45 · outbound

This paper cites Black box variational inference.

Preference learning made easy: Everything should be understood through win rate Black box variational inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.507183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.183300Z digest=sha256:8ab5d623dd1c4cac96863b3d5776e0aacd5cd36ccab5469c2077a00b26d68dba

Observation 2b6830b4-152f-418b-afd5-34dffd1e055f · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

Preference learning made easy: Everything should be understood through win rate Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.189249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.189249Z digest=sha256:9fc2d69084536e27ad7041f50751d37cd88c35abb5e079cb64e57184db2c19cc

Observation a389f205-5390-4018-9842-8d8291d01b0c · outbound

This paper cites Vanishing gradients in reinforcement finetuning of language models.

Preference learning made easy: Everything should be understood through win rate Vanishing gradients in reinforcement finetuning of language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.481360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.195177Z digest=sha256:b565e092ebdbb249c91c77fd3c97a20e5fe7b242613c1f5e3153d63f0ee82dfe

Observation cacef77f-dca8-49dd-a1ec-72cd0c5e9daf · outbound

This paper cites Direct nash optimization: Teaching language models to self-improve with general preferences, 2024.

Preference learning made easy: Everything should be understood through win rate Direct nash optimization: Teaching language models to self-improve with general preferences, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.452006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.200061Z digest=sha256:b62167746b333399c1488294e278f3fa968db6afe27615882f0d9566cbd0328a

Observation 64519d34-4c42-4d2f-a260-b89f02cfed61 · outbound

This paper cites Direct Alignment with Heterogeneous Preferences.

Preference learning made easy: Everything should be understood through win rate Direct Alignment with Heterogeneous Preferences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.206355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.206355Z digest=sha256:8b1d476d63a381d74faec57b29edde09fd47110f3c6ffd6c0b35f70a46ada8eb

Observation 823f03aa-ab4a-4e29-9d59-69a72139f1aa · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Preference learning made easy: Everything should be understood through win rate A Long Way to Go: Investigating Length Correlations in RLHF

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.215175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.215175Z digest=sha256:6c32a3bdfa425571b3fa3d49c57c8b07ffecca85173b3a830b8ae1ad3a7311e8

Observation fea3bb9f-372c-4b55-89ef-4ac8e716d64b · outbound

This paper cites Distributional preference learning: Understanding and accounting for hidden context in rlhf.

Preference learning made easy: Everything should be understood through win rate Distributional preference learning: Understanding and accounting for hidden context in rlhf

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.424473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.221926Z digest=sha256:c1b8a1ca804f9b361db6c34138249adcc8e9406061548f2f88bade0e2c3e2719

Observation d577a9f7-535e-42c8-8c1a-7fde56c8e413 · outbound

This paper cites How to compare different loss functions and their risks.

Preference learning made easy: Everything should be understood through win rate How to compare different loss functions and their risks

Reference 46

Resolution
verified exact
doi, observed 2026-08-07T18:32:21.406280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.227381Z digest=sha256:59e136caebf33a85eeffe4b2b0b8c7bfdcb59b7232e0283d7dbcd73838fdcfb3

Observation 3b65fc6e-bc9e-44f5-a051-1c4c21f62af3 · outbound

This paper cites Learning to summarize from human feedback.

Preference learning made easy: Everything should be understood through win rate Learning to summarize from human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.232480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.232480Z digest=sha256:5937ae5b4912fb483309d9b4ed5154bb31608025595aa5d676ca1b87a20c3a5c

Observation a9667a6d-961a-4104-b483-ad8b2d1fbb5f · outbound

This paper cites S., and Agarwal, A.

Preference learning made easy: Everything should be understood through win rate S., and Agarwal, A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.394366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.237661Z digest=sha256:e9e95b9b758687fe5faab41f9ca9ad4a04d09250cf2cdd89cce14d169d0d8342

Observation 59644008-f9e1-4566-9f7d-bb638ceb40cd · outbound

This paper cites Preference fine-tuning of llms should leverage suboptimal, on-policy data.

Preference learning made easy: Everything should be understood through win rate Preference fine-tuning of llms should leverage suboptimal, on-policy data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.375130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.243309Z digest=sha256:93059a4bc28dbe2c3d52a7ebcb98a43931ca40f17e0ff148f5a3977c6c716e5a

Observation a8b1c143-4b83-49ea-bf47-4462a0fc4edb · outbound

This paper cites D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P.

Preference learning made easy: Everything should be understood through win rate D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.348188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.250115Z digest=sha256:e70d7197c07eb7cfb7e608b6343e191ab5fd8fcde1df75439f5348fe9d456c58

Observation 8a4e59a3-d701-4d99-98be-dd44878199ec · outbound

This paper cites Trl: Transformer reinforcement learning.

Preference learning made easy: Everything should be understood through win rate Trl: Transformer reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.322709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.259111Z digest=sha256:22f32c00aa761fa8de70a38430cbe03637be0a07241f5b0e52151365c373b12d

Observation 14b72f64-ca56-407f-aee4-3b0792b5404b · outbound

This paper cites Enabling language models to implicitly learn self-improvement.

Preference learning made easy: Everything should be understood through win rate Enabling language models to implicitly learn self-improvement

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.297854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.265704Z digest=sha256:1b2ecf46ab542e26369107e7a04c1c6d5cf201f11297ae2827e6f9c721f67acf

Observation 622b1c22-0d0e-4116-9b0c-ca6b54aefe74 · outbound

This paper cites Transforming and combining rewards for aligning large language models.

Preference learning made easy: Everything should be understood through win rate Transforming and combining rewards for aligning large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.261958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.273577Z digest=sha256:1e012c2f0bb0bcf8b78c7a561f80103d3f6729c6200e099d2e2620b9c66b4711

Observation 14cf1242-2bb1-4135-80d9-aa4b0aefe821 · outbound

This paper cites Policy gradient algorithms.

Preference learning made easy: Everything should be understood through win rate Policy gradient algorithms

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.238493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.282655Z digest=sha256:704440d5f26ec96d899c742717a9dd53adeb214065c5c52e7eca23e35fa3a428

Observation 74368c74-59ad-45f1-a388-e08c26b885b0 · outbound

This paper cites V., Murray, K., and Kim, Y.

Preference learning made easy: Everything should be understood through win rate V., Murray, K., and Kim, Y

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.201317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.289186Z digest=sha256:c802f95dab2998e71ce58d7e58f6f5f030ca542898a1dc2fc044d603bc2c2cb3

Observation 97ad70a4-9085-4d57-ad7e-f3d6738836c3 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Preference learning made easy: Everything should be understood through win rate Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.300233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.300233Z digest=sha256:d5ee4b40a83521492e771fac633ce487892e68bc2864e8423bba5e5cbeba426f

Observation 00f98043-ba0a-4e26-a390-a345ade3e382 · outbound

This paper cites Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J.

Preference learning made easy: Everything should be understood through win rate Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.175500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.307867Z digest=sha256:186c311bc97df048cbc8da713063fcd723e1326881973664f24632a8f4935437

Observation 5e8863dd-e6f4-4a5c-8ebe-e6fde31070dd · outbound

This paper cites an unresolved cited work.

Preference learning made easy: Everything should be understood through win rate Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:32:22.143704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.312919Z digest=sha256:282e7c1343738155c7466ca9c0b7910a075545c7f70c6acdc2e07bee691e3c3d

Observation 862eb67f-a726-4035-9669-b97892d90489 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Preference learning made easy: Everything should be understood through win rate Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.120972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.318779Z digest=sha256:9216f82b2e4607ad3f0de35deea894c1c9dff1ac7a6a8520f67d1325d4200e5f

Observation e9610168-ebbd-4b38-b0c3-8b7fab8f24bc · outbound

This paper cites I., and Jiao, J.

Preference learning made easy: Everything should be understood through win rate I., and Jiao, J

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.100251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.325381Z digest=sha256:eaa729a897ed16067551cccc519e62ce2a66381c57322a0f538b29d5e94e2e41

Observation 4419939a-a0f5-49a7-b640-ecc4d607b2e2 · outbound

This paper cites write newline.

Preference learning made easy: Everything should be understood through win rate write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.331060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.331060Z digest=sha256:bd6b01ed665df98d9b691a0609194b04cb1dff614b9420c3e3cbc3b8ffa760d3

Pith citing papers

No inbound Pith citation observations are available.