Pith. sign in

Paper Citation Record · LEDGER

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.07725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07725 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:38:10.381600Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 757181e0-3525-40f8-bb06-84aa17af40a3 · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.762128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.762128Z digest=sha256:c6efedd118f14f6519d8591bc7e5e7ad8173cd76aaf3cae37275e880e67ecb95

Observation 96c56c30-3ddb-42ec-a8d2-3f4ce638959a · outbound

This paper cites the method of paired comparisons.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization the method of paired comparisons

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:38:10.807571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:38:09.793593Z digest=sha256:dd7e28388f2fc8e48c944d9d3e8155a893dc2db794324149237aaa49811a493d

Observation 2a3dfb02-53da-46da-a6b8-b9cb4713ac29 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.822571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.822571Z digest=sha256:67d8468ec8a126a3cb9a6594c12da93077b3cc9f4a3dee2b6c5a0b62b3ae01ec

Observation edf792c8-05a7-415b-8318-369e1b8a0856 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.880112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.880112Z digest=sha256:955e70dd62f26eb38c448a61e5902983f1a9f4f920d439b61b43b5a495ee1f64

Observation 9f0bc812-d474-4d61-bbfa-f4c4e9543a29 · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Direct Preference Knowledge Distillation for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.954900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.954900Z digest=sha256:cb6b5283c559ccf1b4ba62852f99cb7e167f380baa2e15cb15b692b6ddd4f085

Observation 0f567fc3-240d-42c3-9de7-06984d22f8d7 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Rho-1: Not All Tokens Are What You Need

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.027182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.027182Z digest=sha256:44f52073d86bd5b353c4d0f0cecb4806ffd170287ea412143d85a791eff93157

Observation b0468e90-10fe-4539-9ed8-a4cd3ff64abb · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.073796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.073796Z digest=sha256:4e8c3d318cd149414e0d1a9747222846d11e0612ae15a0059d178f8e7d9f46f2

Observation 3fd08862-22af-4123-86d5-2c1132d1fd50 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.166687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.166687Z digest=sha256:ef18078f01ba0aee8332f72e94bb9fbf395df7e9e065503435065ac0bf1cd2ed

Observation 542a403c-c130-4b1b-b348-5626121a0de7 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.193064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.193064Z digest=sha256:dbe1216d3b172ad02edba388a94a3f0a0c1665b31c3bc0df4358ceead6870a15

Observation 1e5aa232-9d66-42ed-861c-1a5db1eb57e1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.241141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.241141Z digest=sha256:b936681458158717f0930661da63df701312c3f4e6b29e37db211d487e4d0672

Observation e2a4ea7f-6a3c-4fa5-9f01-286a080192e0 · outbound

This paper cites an unresolved cited work.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Unresolved cited work

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:38:10.624697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:38:10.308061Z digest=sha256:ed8c015695aee70db9b3daa7aa35e95a6e129ffceb40cbdeb226388840aeb40b

Observation 2aaae1ee-68f1-41d5-b2ac-a08da058766e · outbound

This paper cites , " * write output.state after.block = add.period write.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization , " * write output.state after.block = add.period write

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.339349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.339349Z digest=sha256:d86ebd22c647985b51935e421e0d71425c9d3cf5c074c24fb80b8032b16ab669

Observation 252f0a5b-4b4f-4d3e-a159-bab57d9c5aef · outbound

This paper cites write newline.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization write newline

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.381600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.381600Z digest=sha256:1d17930000801916257fdb9319a72480b6c55129ba58fc8590e56e3630313f64

Pith citing papers

No inbound Pith citation observations are available.