Pith. sign in

Paper Citation Record · LEDGER

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2505.02363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02363 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:00:59.166273Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afa7c4fe-c646-41fd-a6bd-0cba96c303e7 · outbound

This paper cites Cited on pages 5 and 16.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on pages 5 and 16

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.366877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.119432Z digest=sha256:49f037a1a2751e2d09b4c272afdabb44443f77cb8c838a6dd1ad7b8cd9c1f311

Observation 34a21b43-a5f4-4ce6-86bf-5441e713869f · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Direct Language Model Alignment from Online AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.123289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.123289Z digest=sha256:57d7947a31d030275aa2bb1d0d44560f3b6f2e607cab06d3e23f1e12b86c7573

Observation cf4e34e0-4cb4-4431-a782-b4a8b9b6f72a · outbound

This paper cites A Systematic Examination of Preference Learning through the Lens of Instruction-Following.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning A Systematic Examination of Preference Learning through the Lens of Instruction-Following

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T01:00:59.264310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.132394Z digest=sha256:16b62177960a64015a214d682521a991a29c87993f0bac82644bc36bd79c777d

Observation b31bec9f-a623-4816-8e5b-1a24272994c4 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.137081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.137081Z digest=sha256:c1141e2d79dcd76531aab2a4f1c95485048995887a37c564769ce889e59ace0f

Observation 793b61c2-6c27-4657-a4c6-d876e1558bd5 · outbound

This paper cites Cited on page 7.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on page 7

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.354235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.141276Z digest=sha256:6e954ed3ee9385c6fa048480898ac1835f061730d92b1f11380d5691e4d823d8

Observation 576fcff5-0de5-47ff-8f57-a900f2251c2b · outbound

This paper cites The Importance of Online Data: Understanding Preference Fine-tuning via Coverage.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning The Importance of Online Data: Understanding Preference Fine-tuning via Coverage

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.145137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.145137Z digest=sha256:a90fe73daf5d1637d8d449130064a41781dc7809a52fdb602613e47b00e40f3c

Observation b00d6e0f-6d90-4bd3-9caf-f5d435a1b991 · outbound

This paper cites cc/paper_files/paper/2020/file/ 1f89885d556929e98d3ef9b86448f951-Paper.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning cc/paper_files/paper/2020/file/ 1f89885d556929e98d3ef9b86448f951-Paper

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.341017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.149804Z digest=sha256:7419f455fe8d188fc430b0d9057010f84db1c1ec2cdca4ee1f070679c36d07fa

Observation 574c9829-6181-4031-af24-46290e623a13 · outbound

This paper cites Cited on pages 1 and 7.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on pages 1 and 7

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.328588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.153805Z digest=sha256:e685d16b2154c6d927a5e638b0035a0d64a099ea27ec48e9288217bcf456f15f

Observation b61b3eeb-439c-4c70-9546-e2dddc94c7d9 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Understanding the performance gap between online and offline alignment algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.162363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.162363Z digest=sha256:bec8d61265fcc526fefd6fe0c723f8b66a2822039158049e482cf01e9370f97b

Observation a1acd6bc-b786-4519-97db-86183bb6da36 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-16T01:00:59.166273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.166273Z digest=sha256:7dbe8692e8483fe5f0a8531e1d593355c8ad753b3749f4acf4a56a027defa9d8

Observation b5cbbd16-688e-4e4a-9e57-33d7c7e4f2ae · outbound

This paper cites Mistral 7B.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Mistral 7B

Reference 309

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.127796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.127796Z digest=sha256:1ffc2f6e7d6b9d1f9a36168e40c45d7f328322d55311fd0d7731964e5c583554

Observation 8e809b21-aa7c-4dbd-9dea-3d88109d9200 · outbound

This paper cites acl-long.662.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning acl-long.662

Reference 662

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.391858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.101760Z digest=sha256:66a46318ffbec88b83a4d2fe9b5f9615934733750b85102d8ed6a873a16cc82e

Observation 041082b1-e0c6-4775-9b78-516ac6f1b0af · outbound

This paper cites URL https: //aclanthology.org/N19-1421.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning URL https: //aclanthology.org/N19-1421

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.157860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.157860Z digest=sha256:2cc5b6e099a7e98e7ec26f243ed7a35de061991154bc64d19160326e47f6260c

Observation 99eb5fc0-a8a6-4d90-b54b-f04130792ae2 · outbound

This paper cites Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration

Reference 2020

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T01:00:59.305430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.114689Z digest=sha256:ade0aae76975e8a17e91c3c5b36cd4934068f7337fddab5972db3442bcc53803

Observation 1fb577bf-7109-4d52-a626-72cd5079c479 · outbound

This paper cites Cited on page 7.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on page 7

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:00:59.379338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T01:00:59.110420Z digest=sha256:9d7b7a23a40b117f4b307688d5ac2d3447293ea1c2947af0ce5f4ea98f9e92d4

Observation e2ed02f9-79ae-42fa-a24a-62ceb5f3e2ac · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:59.106363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:59.106363Z digest=sha256:a83225f231219acd57e6170ce67d98fc821cc6913fbe7368a4667272c701c99e

Pith citing papers

No inbound Pith citation observations are available.