Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:41:20.743435Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2412.15282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:41:20.743435Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 44a89a05-35e4-4137-b593-d08edd9d19ed · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following The Llama 3 Herd of Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49292da-55ef-49b5-96b3-85483da3f51f · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d803d4db-5551-45f4-a5c0-be8fa0309680 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c7ddc2-4d30-4df0-a18e-36bbb5d2d883 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ad7b07-5d02-43ae-a588-8ce714176bf7 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6784637-1a2b-4c90-840f-60079620f06e · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa0feaca-9e16-42a4-91b4-f2f712e84369 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following GPT-4 Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f482485d-3b53-4084-a950-40d7f860e247 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following GPT-4 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af80b31-0e51-4f03-909c-6d85ce533158 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Iterative Reasoning Preference Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c20e5985-e0d4-442f-81a9-344753191cdf · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Iterative Reasoning Preference Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e61e3d9-d6d9-48f1-9f7e-b6369487c1ab · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64414039-2874-4c9c-8f3c-5600256de41e · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation afd7252f-1140-4724-b201-b1ce7948a376 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Benchmarking Complex Instruction-Following with Multiple Constraints Composition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95259ad0-a5c4-45cc-bcae-a6225943434e · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Benchmarking Complex Instruction-Following with Multiple Constraints Composition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f028a299-d28d-459f-84e3-0fa6a25f4242 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e51b3f5-5634-463f-bc0e-2fa5f74376bd · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a3b165-059b-4ef6-ae1c-2697585c1088 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da23cc91-d9c0-4101-9133-778167ea3036 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7194f92f-67ab-42f1-8ad8-6d1761891809 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Instruction-Following Evaluation for Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca6d087-992e-4ce6-af47-1f88ef4048e8 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b6813e-b368-44e8-b7c9-e9c86e13b50f · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a785059-adbd-411d-8285-ae033ae23db1 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d7dd16cd-5fd8-40b3-9895-2adc459dd902 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93cfca34-28bb-4a97-a4ac-f5d86d78c3f2 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c8ccd28-1f4c-40b4-aadd-3cd7f07c2584 · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Gemini: A Family of Highly Capable Multimodal Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0c2472-2aa9-484d-ae25-8e4a897602eb · outbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following The Llama 3 Herd of Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.