Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T01:00:59.166273Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2505.02363.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T01:00:59.166273Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation afa7c4fe-c646-41fd-a6bd-0cba96c303e7 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on pages 5 and 16
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 34a21b43-a5f4-4ce6-86bf-5441e713869f · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Direct Language Model Alignment from Online AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4e34e0-4cb4-4431-a782-b4a8b9b6f72a · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning A Systematic Examination of Preference Learning through the Lens of Instruction-Following
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b31bec9f-a623-4816-8e5b-1a24272994c4 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793b61c2-6c27-4657-a4c6-d876e1558bd5 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on page 7
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 576fcff5-0de5-47ff-8f57-a900f2251c2b · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00d6e0f-6d90-4bd3-9caf-f5d435a1b991 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning cc/paper_files/paper/2020/file/ 1f89885d556929e98d3ef9b86448f951-Paper
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 574c9829-6181-4031-af24-46290e623a13 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on pages 1 and 7
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b61b3eeb-439c-4c70-9546-e2dddc94c7d9 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Understanding the performance gap between online and offline alignment algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1acd6bc-b786-4519-97db-86183bb6da36 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5cbbd16-688e-4e4a-9e57-33d7c7e4f2ae · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Mistral 7B
Reference 309
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e809b21-aa7c-4dbd-9dea-3d88109d9200 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning acl-long.662
Reference 662
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 041082b1-e0c6-4775-9b78-516ac6f1b0af · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning URL https: //aclanthology.org/N19-1421
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99eb5fc0-a8a6-4d90-b54b-f04130792ae2 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1fb577bf-7109-4d52-a626-72cd5079c479 · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Cited on page 7
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e2ed02f9-79ae-42fa-a24a-62ceb5f3e2ac · outbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.