Pith. sign in

Paper Citation Record · LEDGER

Debiasing Online Preference Learning via Preference Feature Preservation

As of 20 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.11098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11098 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:32.728004Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97cd3f7b-566b-455b-9cdd-a3f0b49084f3 · outbound

This paper cites online" 'onlinestring :=.

Debiasing Online Preference Learning via Preference Feature Preservation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.520994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.520994Z digest=sha256:6e8571bf61b99faefb6aa9c9dfa77158a72003e96e314ee04fd019b117c7b2ab

Observation c393ac49-6827-468f-9471-6b4e1b397011 · outbound

This paper cites write newline.

Debiasing Online Preference Learning via Preference Feature Preservation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.526128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.526128Z digest=sha256:dc66b8c3ef34eb941bf3c642afcbfe5eb53b12662c815d9801d904652d3bf4df

Observation b14c64b4-b5fe-4978-a3d5-703f650aecdc · outbound

This paper cites write newline.

Debiasing Online Preference Learning via Preference Feature Preservation write newline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.531806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.531806Z digest=sha256:ab57493e7d8d09c9cc493de1dcdf919dc1aac1db27b8b8e1a0209bc90b489e66

Observation a67aad3a-e58a-4e61-a615-7a741689fdc0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.532490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.536525Z digest=sha256:080e021ccd2e55663d171293416195f05f695349d0a64b6f60788e295f677862

Observation 2acc3515-d81f-4cc2-8512-0b46352bdb5c · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.518687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.540968Z digest=sha256:8c60ac077ce828d7526eef38de8b5999d11b0236188bfada6d5e17264e57c29a

Observation 44a430db-54fa-440b-b31f-c752a6d5620a · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation A General Language Assistant as a Laboratory for Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.545630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.545630Z digest=sha256:5a6bbcd61766868014605e558018d2ff6fb53b0d59a5054d435ff7605ad7efcb

Observation 54f27b40-29ad-4e25-8375-2ab3cc997f5a · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.504015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.550254Z digest=sha256:0a719f4ff3ffa56b84c957317f8b202e8e2790b802126cb4bddf0ccc40fb645a

Observation 4638ed7e-9013-4d13-a010-f0d005282a97 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.554671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.554671Z digest=sha256:16a6aeb503bba83140c459a7229a0a1ba6e8cd98f23854cf681856f63ea28085

Observation 1fa1e298-330f-4b76-91a0-94a7d1af9242 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.479058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.558585Z digest=sha256:2f63e8166dce2c0fdaafd357e2bdb3f2ee1d5598c1dc5f2c41e4cf6cb3c94879

Observation 39cd5008-b9f1-45cb-b3e0-b8c2368a4106 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.466212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.563093Z digest=sha256:a4c4be183aa91e9d36f236145609767b65ffdb38eab74d7dc4f568bc04e44c04

Observation 15bfb751-8677-430a-ba9e-c97906086dd8 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Debiasing Online Preference Learning via Preference Feature Preservation UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.567554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.567554Z digest=sha256:2f133d0b0b619ba7f645924bc9b056bcc6ef8894b89c6817d79b74f78958fa39

Observation 37b32f76-af8f-48b6-844a-3ffdeb628b8b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.452443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.572066Z digest=sha256:1967e25b19fc89b4d1c30c39eb22e95eac81ab50f67c6fd9ca1c7208d37189ce

Observation 14915f7f-eb95-403e-91cc-6af505725a08 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Debiasing Online Preference Learning via Preference Feature Preservation Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.576314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.576314Z digest=sha256:673895512e4da87d0d168036d07bacc6f087d4006b63881175678c06ee208e76

Observation bbe05557-40df-4a4b-a544-a963a7750947 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Debiasing Online Preference Learning via Preference Feature Preservation RLHF Workflow: From Reward Modeling to Online RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.580792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.580792Z digest=sha256:27f57909664808ae0a26449ba0b7e6401c2bba15d0fcf915d3322ef460413e63

Observation 67bf7084-0174-4fcb-99f5-a281574c56aa · outbound

This paper cites The Llama 3 Herd of Models.

Debiasing Online Preference Learning via Preference Feature Preservation The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.585102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.585102Z digest=sha256:ada181d970c0bed99b39c7d27820badc0b5e0273b6a27dd2131230f7ba342d70

Observation 0209bffc-0395-4431-9978-bde3421de69e · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Debiasing Online Preference Learning via Preference Feature Preservation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.589245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.589245Z digest=sha256:db2cdc44243fa0861c65b6e24e74d54befdbf89d5052c3b579801136685b2362

Observation ec8188f9-2433-4f22-a2bf-cdf944df81a2 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.438711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.593625Z digest=sha256:86803fb279ffd63bfd0e512d437330c2eff78198260ac2edaf8541db134a38b4

Observation 4a59ccb9-e276-465a-8636-d63bbbbe37c0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.425259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.598590Z digest=sha256:73484785a23ce7113efe1d0ffc377af0882df03077b3e953bfad9aac77c6fd1e

Observation af5dfd71-9369-4733-ac9b-992386de1f21 · outbound

This paper cites Mistral 7B.

Debiasing Online Preference Learning via Preference Feature Preservation Mistral 7B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.603244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.603244Z digest=sha256:e0d0b99c673dcc6addcef571c9598053b5f1363d21e30818ac8840e999a62bd7

Observation df7810e6-6911-4b53-9744-df9bc024dc31 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.412157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.607485Z digest=sha256:6ad7fa303172ce125702e43c01475fb8c0b7425504a5af1bc866170a3fce2824

Observation 231ed846-a0a5-4267-8112-b3975ad68ce0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.399238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.611469Z digest=sha256:8ea32dcfe3f7d0cecbeb1e21a527679b9537eb14d29365f1db2f0cd0750bd1e6

Observation b44231d1-45fd-4c50-8cbb-88ba9410d503 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.385631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.615612Z digest=sha256:76ac408973e532c19d8f15a537a9f7bfcb18c9e4c776abc44aee416a4cec31bb

Observation 2c603c55-2ae5-4eba-ab22-10ea4b4bced8 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.372321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.619403Z digest=sha256:d06c71f24a86cc476801b0ff3da94cc0b1ba214f86f68080713baae4b4279728

Observation 3955dbab-d44d-4b18-82f6-e7a3f4bea961 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.358683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.623578Z digest=sha256:2854c825ec5f2a328dac7a7c6d8f75ec568ca19fb1a106ade78bc3db76e4a495

Observation 94169d60-021c-4f72-8000-9e67f334bd5d · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.343913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.627509Z digest=sha256:2c3eb2fdc166aaf95f5bf81aeea0808c3f247590da52c6a2b43178c6c1c9df73

Observation 7c27b1ca-6f27-474e-8f27-021ee75597fa · outbound

This paper cites Dissecting Human and LLM Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Dissecting Human and LLM Preferences

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.631391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.631391Z digest=sha256:b10158aa0321f2e1455a84729f2098f6bffa207304b797d08200fce5116bce65

Observation 4cd640aa-6fef-45cd-ab5a-98e22d37fa91 · outbound

This paper cites Decoupled Weight Decay Regularization.

Debiasing Online Preference Learning via Preference Feature Preservation Decoupled Weight Decay Regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.635828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.635828Z digest=sha256:a06f94adaf584df548473b90e09cee3244aa954582abbfbce56f705cb845f1e6

Observation 6c8ed455-6fc1-4c28-9a79-47d7443ba863 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.330752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.639794Z digest=sha256:864066bb54401bff93496920c52a70fe442773c66bf4c740b9819630f9d7135d

Observation 0050e4e5-2c19-42df-b80a-5aefc823c192 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 29

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:05:33.018292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.643899Z digest=sha256:6f1a2e6ee187eb97d44060cf27b18b814922eae5236e6ba5a17f0863099ef3c4

Observation 134dfb4e-b7b2-4554-a56a-0ec9ccb4fb01 · outbound

This paper cites GPT-4 Technical Report.

Debiasing Online Preference Learning via Preference Feature Preservation GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.647804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.647804Z digest=sha256:4af13f2fbb7efce3bba44cd87e2824f2d62693a5b3ad1d270b719fff4e13316e

Observation 0d800e66-12e2-4516-a0b7-70f26436693b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.316887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.651800Z digest=sha256:7436b92be0f2ff0a124fe26fdf39cf5d677f148917bc98cbc36e99514becbb04

Observation c16bedea-7899-4537-b161-7e33e1cf8054 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.303371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.656381Z digest=sha256:1ad92b44b8160d783e55e73ca33985468a6ffb97589e88f83894017295e5e4fa

Observation e3c64ba8-a585-434f-8d09-81744b4e0b12 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.289403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.660320Z digest=sha256:32e44fe6aacfe06e9adea75b144f8a4c3b90018ba83d6e3f456815e6aec6af4f

Observation 7ccfae8f-eb6d-4547-99b4-4b9439e73479 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.275333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.664289Z digest=sha256:a5fd917b2256d44e9a2b407d25a6dea649d38f87a74f94728c60a9d23846f9fb

Observation 0449bcc5-cd22-43c8-8a98-5a9be9064e05 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.262059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.668431Z digest=sha256:e45dc89e98a86e85cfe955f9470cfbeb85f6371ab5dcc4b8a4f59e25da17b4ae

Observation 0fc2bec0-5495-40c1-a5e9-18cfb280fb3c · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.672178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.672178Z digest=sha256:eae883e177de4b4a5703e5d14a72c6df98616175697cff03504fdf9f6ff97038

Observation a3844b5c-fc83-4018-b6a9-5f94054c4f26 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Debiasing Online Preference Learning via Preference Feature Preservation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.676825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.676825Z digest=sha256:0feb4dd4ec8440790f4d6504567a54d8a4b4e69b10ea35fe6d6d60c467448ac5

Observation c1ca1cb8-b1c0-4130-afcc-3099e44ece02 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.248054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.680958Z digest=sha256:a8f3119fc964566c8fbb609e9bc633cb3e6a784d8ffd251947b0de08fbc912d5

Observation 0b4f36e2-831a-4a94-b610-5e9d3d917d0f · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.232005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.685213Z digest=sha256:e6e513aff710c3ec6af151a7be19fc10edb1b5a3bcccda68c15742cadf7a7fd3

Observation 362a54b7-755f-4ab6-b9d8-4a615b1153db · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Debiasing Online Preference Learning via Preference Feature Preservation Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.689590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.689590Z digest=sha256:d1e59db7cac18751416bd7bb69499ae92a9165c1221b87e493eb928a746726f0

Observation 242ee8b0-5754-4c86-9e0e-6a9315e892ae · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation Zephyr: Direct Distillation of LM Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.693883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.693883Z digest=sha256:34f27101b279dc03cdf07f84fbc98bb443a77b21e034705fae324bc301f9e6d8

Observation fdd9d55c-10bc-41fe-9be2-aeac5ce84b3b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.218139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.698505Z digest=sha256:ec9bd3a23afad6fe6b2d83fee28df5ef769c22b7fce9e89c011e3e28b757d9c0

Observation 9192bb21-5105-4608-9b76-1f00beb25dc1 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.202246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.703255Z digest=sha256:8e5ccb9b7ea79ada933ada21ce9bc4c9b5e07bf4ebebc0c8072508f6d295210c

Observation 0d0b9669-c8b2-4b83-8629-aa35d4e4397f · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation Self-Play Preference Optimization for Language Model Alignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.707537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.707537Z digest=sha256:2fe1a0d8cb97ddf715a193036c0021dacb06a425b891edf21d1b0d6723289fb2

Observation 944ea925-9df0-422d-bede-ecfb6cef2c23 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.188133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.711756Z digest=sha256:c0198d8c40d13037305f3e1934e542fef4014e7245f0b0e8d84cc1b967ea2396

Observation 8a00e02c-ef7c-4cd6-b98f-5b30ad389630 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Debiasing Online Preference Learning via Preference Feature Preservation Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.715738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.715738Z digest=sha256:c06b34b06268b52edf9105da0601d07f772850f238cf2b8e339a403218ddf2ff

Observation 9520ea74-8889-43c0-bd1b-372e99a61b44 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Debiasing Online Preference Learning via Preference Feature Preservation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.720004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.720004Z digest=sha256:9959af126ac64e2dd2548722ba701f40a5ae2d4177c070b8f18658faf1021dc4

Observation 0bea02bf-8970-4335-b3c0-29668a3e6023 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.174460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.724090Z digest=sha256:7955a65e2e25355cadb114aae8c98e0fe7f7907f936101b0a17bc7bcf505e81b

Observation 385a7655-4429-4b52-86ac-4b787c317a94 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Fine-Tuning Language Models from Human Preferences

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.728004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.728004Z digest=sha256:4c80c3f843bc412a567a3f4d6d0276f4a722f2e55d5479db7908b91c45e807a3

Pith citing papers

No inbound Pith citation observations are available.