Pith. sign in

Paper Citation Record · LEDGER

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization

As of 19 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.07725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07725 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:38:10.381600Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 757181e0-3525-40f8-bb06-84aa17af40a3 · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.762128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.762128Z digest=sha256:962d056f5a5fcc364266804f414a2cf7c172d330be77c2a4a76a7156412c4f04

Observation 96c56c30-3ddb-42ec-a8d2-3f4ce638959a · outbound

This paper cites the method of paired comparisons.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization the method of paired comparisons

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:38:10.807571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T18:38:09.793593Z digest=sha256:d958b343e88a540ead7a1218a7d68190710630df644be8b7404b90cd5024fba6

Observation 2a3dfb02-53da-46da-a6b8-b9cb4713ac29 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.822571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.822571Z digest=sha256:0f57cda475719e3de060c1ccb5322cb4f6d958670f7d64cd061a432ae13ab859

Observation edf792c8-05a7-415b-8318-369e1b8a0856 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.880112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.880112Z digest=sha256:78b720ca9c55951989dbd9f458b1c671bb7fc81c1f0c8ab5f77d4e63b6d2bdd3

Observation 9f0bc812-d474-4d61-bbfa-f4c4e9543a29 · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Direct Preference Knowledge Distillation for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.954900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.954900Z digest=sha256:73e03859ec938a418f52f193ad2c4694fefb74d876bccb4a558d78a2fd9b5916

Observation 0f567fc3-240d-42c3-9de7-06984d22f8d7 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Rho-1: Not All Tokens Are What You Need

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.027182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.027182Z digest=sha256:459edb45f51ac9e3f0602e8dfcc575b1a738c2c10552b7b2ca7666f7d7b67f7f

Observation b0468e90-10fe-4539-9ed8-a4cd3ff64abb · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.073796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.073796Z digest=sha256:e97d50b07133d6d444acff83c611c41f240ab8403b45aa9e76edc3ec796bcaa1

Observation 3fd08862-22af-4123-86d5-2c1132d1fd50 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.166687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.166687Z digest=sha256:a81778e06990f1201f79ba421b61f28cea7f4dbc4b4966e331396830fbbd3b77

Observation 542a403c-c130-4b1b-b348-5626121a0de7 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.193064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.193064Z digest=sha256:d16e9f6a5678aba81c8a7759789791d4b605bb564e2ac2da0ad1003911a64cda

Observation 1e5aa232-9d66-42ed-861c-1a5db1eb57e1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.241141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.241141Z digest=sha256:db6ed9c12923992ab59d433f80444e996e4566baa862e3f94ca1e4f8875f852f

Observation e2a4ea7f-6a3c-4fa5-9f01-286a080192e0 · outbound

This paper cites an unresolved cited work.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Unresolved cited work

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:38:10.624697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T18:38:10.308061Z digest=sha256:ab4fbd3782de3492d0d7dd9aa3418ea880a88f84eee795727589e1df23e1d9f1

Observation 2aaae1ee-68f1-41d5-b2ac-a08da058766e · outbound

This paper cites , " * write output.state after.block = add.period write.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization , " * write output.state after.block = add.period write

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.339349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.339349Z digest=sha256:df4f4bdd896d460561422cf4d7131bf701ed4fa966f578a094d1183b56d8e91e

Observation 252f0a5b-4b4f-4d3e-a159-bab57d9c5aef · outbound

This paper cites write newline.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization write newline

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.381600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.381600Z digest=sha256:cb2ba6b43405fcce8259983badc2a77c91f407217c1e2c3859e249ef0559d737

Pith citing papers

No inbound Pith citation observations are available.