Pith. sign in

Paper Citation Record · LEDGER

A Systematic Examination of Preference Learning through the Lens of Instruction-Following

As of 13 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2412.15282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15282 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:41:20.743435Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44a89a05-35e4-4137-b593-d08edd9d19ed · outbound

This paper cites The Llama 3 Herd of Models.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following The Llama 3 Herd of Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.670025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.670025Z digest=sha256:c6439dd90f3aa104852c99e9c486e87737f0dc602b384f7de3a2e9fd3f2c0ae6

Observation e49292da-55ef-49b5-96b3-85483da3f51f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.676408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.676408Z digest=sha256:07409432ed616a184b63b6fa5fa0733c225905bec224eb9125dbe3d1076316de

Observation d803d4db-5551-45f4-a5c0-be8fa0309680 · outbound

This paper cites Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.689532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.689532Z digest=sha256:1ac9966164049139d3207297f8c01c3209ad93e46bd610ed8ef63c5d761da288

Observation f8c7ddc2-4d30-4df0-a18e-36bbb5d2d883 · outbound

This paper cites Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.692264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.692264Z digest=sha256:8b0d5e637f91e752794746deac04ddc98d89f993fb3d07a994a9c65c46dc6318

Observation 37ad7b07-5d02-43ae-a588-8ce714176bf7 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.695300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.695300Z digest=sha256:9dfd1be086e5987d67203a1dd082c7624c51101dfc7096c28d6802cd05013c54

Observation d6784637-1a2b-4c90-840f-60079620f06e · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.698420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.698420Z digest=sha256:891a21da10365c42fdd75e211e19391b3a1281eb11109957e2b40a10cdde5860

Observation aa0feaca-9e16-42a4-91b4-f2f712e84369 · outbound

This paper cites GPT-4 Technical Report.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following GPT-4 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.700856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.700856Z digest=sha256:68907df7446cc4273fb501cf05a8d240abf16c4bd0ca820555984b595e6298d8

Observation f482485d-3b53-4084-a950-40d7f860e247 · outbound

This paper cites GPT-4 Technical Report.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.703382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.703382Z digest=sha256:2b757eb3ea2631157bc71bc33bae2c67cdfdfd4f5438c4bcd8e85cfd07f5187e

Observation 1af80b31-0e51-4f03-909c-6d85ce533158 · outbound

This paper cites Iterative Reasoning Preference Optimization.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Iterative Reasoning Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.706243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.706243Z digest=sha256:67c1d140b9ef613206c8149614e5c6fa1eb8266b5b5c7a3a04235fbf005c4a24

Observation c20e5985-e0d4-442f-81a9-344753191cdf · outbound

This paper cites Iterative Reasoning Preference Optimization.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Iterative Reasoning Preference Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.709377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.709377Z digest=sha256:62bf958f96e8b16e00d23156bb4f71e671582233bfadfe87d3bca9eae27fc3c7

Observation 3e61e3d9-d6d9-48f1-9f7e-b6369487c1ab · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.720234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.720234Z digest=sha256:caacdfe29489020e9494a000138ecd84fe8ecfb10c619c25dcc9175c16afbe58

Observation 64414039-2874-4c9c-8f3c-5600256de41e · outbound

This paper cites Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:41:20.911248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T12:41:20.722980Z digest=sha256:31eec14f128c3daf2b966ce993dc20f12c4a2b83999fd87b5f51bf5b72a7cf9f

Observation afd7252f-1140-4724-b201-b1ce7948a376 · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.725094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.725094Z digest=sha256:a739b8e65f5664c4e0f8ea893297e55a4648b1596c29e471950dba64acbab53d

Observation 95259ad0-a5c4-45cc-bcae-a6225943434e · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.727453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.727453Z digest=sha256:05b0e217cfd9634c3b64bfa7ba51135eca489b89234d13bd41c44f408a437659

Observation f028a299-d28d-459f-84e3-0fa6a25f4242 · outbound

This paper cites A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.729703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.729703Z digest=sha256:a7fc4932d0b520dbe46ed508beea4e72ee141a5ef08bdb8751a59716fc83e2f1

Observation 9e51b3f5-5634-463f-bc0e-2fa5f74376bd · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.732265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.732265Z digest=sha256:c4b944d0e8d97665aa076f9d8ea417f687f267dcdfc8de1d13fb9d73d3b7af9a

Observation 20a3b165-059b-4ef6-ae1c-2697585c1088 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.735095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.735095Z digest=sha256:d8e27a10637a99e7d2ad0b9a55a1bfc5aa4e683705fb99eb736c7ac814ddfc06

Observation da23cc91-d9c0-4101-9133-778167ea3036 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.737718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.737718Z digest=sha256:d64c924efd45d93fc54c579f1a02813e4981db2992da3c50739ea0847a137943

Observation 7194f92f-67ab-42f1-8ad8-6d1761891809 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Instruction-Following Evaluation for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.740628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.740628Z digest=sha256:0f0b7633c95cac66c7e240eb1619ebac0a22354234736844412c30ed6e3ea9f0

Observation bca6d087-992e-4ce6-af47-1f88ef4048e8 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.743435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.743435Z digest=sha256:4b3283a10353051a02746ff298c08b9e1cdf1779d7ae29ba5f9880bcdb41a18d

Observation 31b6813e-b368-44e8-b7c9-e9c86e13b50f · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.715053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.715053Z digest=sha256:43d229d5fc0fb97f8324c95880190ec0bba7ea3d20d5ade242269a300eb07909

Observation 7a785059-adbd-411d-8285-ae033ae23db1 · outbound

This paper cites Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:41:20.920072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T12:41:20.717800Z digest=sha256:307e88257a202a14c599b93ad26044cdffb101b4b655cb674cc5263a9ddd5b84

Observation d7dd16cd-5fd8-40b3-9895-2adc459dd902 · outbound

This paper cites Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.686480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.686480Z digest=sha256:012008a6be32a10731b373d784e2cbfb47ea647eb704061da0d10016e5f322af

Observation 93cfca34-28bb-4a97-a4ac-f5d86d78c3f2 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.680428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.680428Z digest=sha256:59430b77ada2dac92706ea52d2d0eb7d374141f697eb0c641dba127b3e2a0f59

Observation 3c8ccd28-1f4c-40b4-aadd-3cd7f07c2584 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Gemini: A Family of Highly Capable Multimodal Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.683556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.683556Z digest=sha256:6ee539315fa56c1720fbd40e9a8762965c77a56527449dde26fec2af78e717fd

Observation 0e0c2472-2aa9-484d-ae25-8e4a897602eb · outbound

This paper cites The Llama 3 Herd of Models.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.673423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.673423Z digest=sha256:756ee65d02100f75ffcdb96cce71228b9b36b26f50c2c578b51da01925153aac

Pith citing papers

No inbound Pith citation observations are available.