Pith. sign in

Paper Citation Record · LEDGER

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

As of 11 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2506.08712.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08712 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:14:45.567616Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:01:59.503415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T14:02:39.828118Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1792dcec-7e4b-485d-aab6-2652c6154666 · outbound

This paper cites Explaining individual predictions when features are dependent: More accurate approximations to shapley values.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Explaining individual predictions when features are dependent: More accurate approximations to shapley values

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.444061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:41.625861Z digest=sha256:b5bc7dbb7c7206c29cb3f071be957d2c024605d9db10a0b943b7cd07fcea01c4

Observation 98329f87-3bfb-457c-ba45-3f3c6e1eacf2 · outbound

This paper cites Llama 3 model card.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Llama 3 model card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:41.672739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:41.672739Z digest=sha256:b1f874e1f6ff38d966ed944126379e463ef5e4082f6f6aafbab2594166cde937

Observation 425b234e-dc88-4f52-914a-71c22079b1b3 · outbound

This paper cites Mitigating reward over-optimization in direct alignment algorithms with adaptive importance sampling, 2025.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Mitigating reward over-optimization in direct alignment algorithms with adaptive importance sampling, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.431146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:41.768753Z digest=sha256:ede2a1219c232356a8f78436f8b2357d06445a602f603a36c05938776101ea8b

Observation 1638c1d0-5da9-4405-87c8-ac49348bd144 · outbound

This paper cites G., Guo, Z.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization G., Guo, Z

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.423282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:41.913775Z digest=sha256:3d151d22e521fd4908b98345f2b82e4afd417aebc15d8594f15658f8b2ea0767

Observation 24dd7f57-28c4-4fe8-9074-5489596df715 · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.023438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.023438Z digest=sha256:5a097bcf468b7f86ced5d0cdb3e59b8d14e83420272508f368e6bd7e99c3bb86

Observation fceb882e-572d-44f2-beca-7305daa69a4c · outbound

This paper cites Step-level value preference optimization for mathematical reasoning.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-level value preference optimization for mathematical reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.135215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.135215Z digest=sha256:9a128d38a35f2dc801c211fb9de6aea514aedfcbb6475841f49c0625f41e5492

Observation eeb05f53-2f33-4bc0-b3c9-4eb5bc98e9c0 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.415696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:42.210513Z digest=sha256:ab84d31edc7703a901266617a7f9f498b8b1f1cf676a3d656cc322cb2655b497

Observation 95b809b0-5d62-4004-b9d9-101bc6a66065 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2023.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Ultrafeedback: Boosting language models with high-quality feedback, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.408306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:42.296984Z digest=sha256:78409a536920441f85eb190d709103c52619c5a685fee129e177116e0bb48bf8

Observation 7e194a62-4662-4361-ba68-22a6d3f5ea1d · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Enhancing chat language models by scaling high-quality instructional conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.398724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.398724Z digest=sha256:4a08cf8b1e0a5926af17548d1e559919a6f2ef8f0ff4888c7d04e14afc1b6383

Observation 45141898-afe8-484c-8e62-c08716403b3e · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.513413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.513413Z digest=sha256:ad15207cef41e0185fb4fe745fb29507798cd0cb768bc9fd8e37adfcd41ab3ce

Observation da680899-0219-4152-ad23-3e93ca90f203 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.597895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.597895Z digest=sha256:d846364f5d739a2690abf92f21f0c954f32bc7a758df4522f505d805cad8fa04

Observation 31236e31-6e00-4b39-85ac-7f5232678988 · outbound

This paper cites Scaling laws for reward model overoptimization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Scaling laws for reward model overoptimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.724379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.724379Z digest=sha256:73ddd341538fdaad7ba5004b44045238910d8d8728b82d3993086b97a1444f48

Observation cda1935b-ddae-4a4b-8f78-97fbe109079d · outbound

This paper cites Beyond imitation: Leveraging fine-grained quality signals for alignment.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Beyond imitation: Leveraging fine-grained quality signals for alignment

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.395535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:42.807169Z digest=sha256:10d6301936c663121e7d26ad2461874c33e6352511862578d60a0b73ef6341a7

Observation a463a060-2d3c-49a0-b719-d5e49148b80e · outbound

This paper cites A probabilistic earley parser as a psycholinguistic model.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization A probabilistic earley parser as a psycholinguistic model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.387379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:42.938224Z digest=sha256:e889da412389efbfe3fe8437e060952afe820611d0d693a391c024b76a0465ba

Observation 4bfa7d60-0423-4c83-aebd-ddfe0126531f · outbound

This paper cites Reference-free monolithic preference optimization with odds ratio.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Reference-free monolithic preference optimization with odds ratio

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.379614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.033275Z digest=sha256:7729ff6eeb7d52097ed09db6a3f8bb67185922843c1887eb72264b2f5682420b

Observation 0350f2ed-d3ce-4d3c-830c-19bf1471817b · outbound

This paper cites and Levy, R.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Levy, R

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.371360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.126624Z digest=sha256:0fba95e0f1da8fb024d8ee4b8b5adcce989fff9fbbf8f0b60d20afeef346db5f

Observation cc52e0c3-8bb3-4d94-a507-dffcc102d1e6 · outbound

This paper cites Mistral 7B.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Mistral 7B

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.201027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.201027Z digest=sha256:19d3012124acf66d923d26b7dae4f3123d47cd709840a9a616991fea73269116

Observation d19fd706-3403-47a2-ad12-cb865ce56a1b · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.248348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.248348Z digest=sha256:a4be2fae1aae8c44f7c35cc314894d1a9b921ada171321cf36c4f5c0ac5ef91a

Observation 20b7bea3-d90f-4df8-87b0-84308a90de37 · outbound

This paper cites E., and Stoica, I.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization E., and Stoica, I

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.363481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.330957Z digest=sha256:99af020b24b86a9f03317a386b87da13398c8481c771d4f7215571a2af3f9e51

Observation 4e3df7f6-7a62-4f25-8a5d-58e625bfe14f · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:14:47.355563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.419011Z digest=sha256:b2763e739cb5c2f301127660166bfb01d044221c66b2bf3489b9c9b3e93a9dbf

Observation ebe34097-d2cb-4aaa-8b22-31ada1e6919a · outbound

This paper cites Not all tokens are what you need for pretraining.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Not all tokens are what you need for pretraining

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.347856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.489684Z digest=sha256:342f4b1ea182e23fc3b0fc07fcdc65e6b66847c40e54a0d8a67bbcf6a8817c1a

Observation 35ed17eb-b8c4-4dc6-a9fe-754b7d2fe017 · outbound

This paper cites Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.340004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.553587Z digest=sha256:35cb2100c603f94122f9d8149809031a966cd663da0b906bce8abd7409581298

Observation c436c405-a559-4406-a993-cd241f3d01e9 · outbound

This paper cites and Hutter, F.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Hutter, F

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.624256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.624256Z digest=sha256:043c42aacc738db2caf65ffd2d8c0d92cf94b0b01ab8abc1e09580786de6eb45

Observation 1adf573f-560b-4750-9223-cce8d4ca61cf · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:14:47.326955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.698453Z digest=sha256:b97eef4f927db40fd1d841fc9f0f0649acf8c04a1311fbcbdb87dda7c30bc5b9

Observation 00841d81-e6b7-49be-92af-ad2c92ce077d · outbound

This paper cites Sim PO : Simple preference optimization with a reference-free reward.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Sim PO : Simple preference optimization with a reference-free reward

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.318519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.768830Z digest=sha256:8a62a3a430c6e5538ce70c3b795faf8a57e68580a5b0ef839e28a5580fc6efaa

Observation 0be6bca5-b5ca-4bc2-8d9a-e9b77595f50f · outbound

This paper cites and Frank, S.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Frank, S

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.851661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.851661Z digest=sha256:ac675bd128dcbf3505285d508455abc1da7f3fd54cd8696bc05eb354e5e1aeda

Observation c3e0bd5a-75e2-468e-b903-a10285da3a45 · outbound

This paper cites F., Leike, J., and Lowe, R.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization F., Leike, J., and Lowe, R

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.309553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.920632Z digest=sha256:b1206472a4c37410c4fac467d79c267c7d0d3b511940beaf9f2b4982630f8f35

Observation 2ee4273c-88b4-4a4d-9b2c-3d9c1bf7385f · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:14:47.262125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.027094Z digest=sha256:406eb39d1039bc04502100f7374c0fe2d9ebe5e013a5f99bffa5af474a690567

Observation 51e5e0f4-f6c6-48d5-9ce1-3c8eade8b949 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.113845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.113845Z digest=sha256:8f49dfb57bd4e4d12565c45e820d89ff0da02395928f2e587f29aecb1a7685ec

Observation 99509864-b239-4a70-b17e-fc59fcc97449 · outbound

This paper cites B., Finn, C., and Niekum, S.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization B., Finn, C., and Niekum, S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.059754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.195859Z digest=sha256:eebd4e24e6a105c8a1b88df945816bc01a7b340cc18cc7b3701f1c273c4e493d

Observation 97ba12ec-4d95-47d8-8976-42f0e7f567b6 · outbound

This paper cites D., Ermon, S., and Finn, C.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization D., Ermon, S., and Finn, C

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.766693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.285480Z digest=sha256:e9bf2050e777d32be72bbfa4b54b84fffce391e920502e8199247e070fd6861c

Observation 84fcbaca-6bd8-4148-bf89-5d8863154608 · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:14:46.666687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.371677Z digest=sha256:5934159a23953d39f85dfc4671341f5377a852c9a803e82f7b2e6e473d481c72

Observation 58b40bde-26e0-4200-923a-e26b358609f2 · outbound

This paper cites S., and Martin, A.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., and Martin, A

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.448597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.448597Z digest=sha256:3096420ed92b3f53450561a74e821c0b9ac9cd10a23605e792642a664abb86aa

Observation 4cbce2f8-8b15-43d9-93b5-5f33853423d1 · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.525429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.525429Z digest=sha256:871f49b66d2d0a5f42e50c88c19ef9b3b3cb232b66cda6a69e0e2c1cdd113e12

Observation db5ae8f2-5af7-4907-8de6-8a392005d2f9 · outbound

This paper cites Trl: Transformer reinforcement learning.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Trl: Transformer reinforcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.606004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.606004Z digest=sha256:63698de6b956be0b27d9f16e002b4bc8adc96ba54dd7ed5774276a93ca2b5740

Observation 5991f4c3-c998-4191-811e-7aca3019706c · outbound

This paper cites V., Murray, K., and Kim, Y.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization V., Murray, K., and Kim, Y

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.544726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.660895Z digest=sha256:a36bff20dff1a8898330eaea9f0ff2f69a4d44fb9a340fbf87fb4eb67edcf919

Observation 8ad94440-41a3-4cd0-9335-cea3803121d6 · outbound

This paper cites Selective preference optimization via token-level reward function estimation.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Selective preference optimization via token-level reward function estimation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.737682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.737682Z digest=sha256:bc4a4b52da12997b4af2caa2f14ee8ace2c8e60648b438373abaa2d73f52ae83

Observation 34656020-91f9-4927-a40e-e1992ba9a49a · outbound

This paper cites S., Hasegawa-Johnson, M.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Hasegawa-Johnson, M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.416795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.802381Z digest=sha256:5274b00ae1936135310582dd5b09e84bd60d685eac362ed4557e240d63f1d309

Observation d628485e-9912-4853-92ab-417c060d13ae · outbound

This paper cites S., Eom, S., Han, G., Nam, D., Jo, D., On, K.-W., Hasegawa-Johnson, M., Kim, S., and Yoo, C.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Eom, S., Han, G., Nam, D., Jo, D., On, K.-W., Hasegawa-Johnson, M., Kim, S., and Yoo, C

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.866870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.866870Z digest=sha256:080b77d2fe07a07c2857b2797dae8b0df8c0210d3f08b4ba51af3c35a1584e21

Observation 683d2268-77fb-4ccf-95d4-895c8b07727f · outbound

This paper cites ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.934080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.934080Z digest=sha256:bd0e3fd8d0ae61a6f585094ce5a9352751de9ab9b201a62ad20ec9bab8cd81d6

Observation a63adae7-2697-41d5-b8ca-70c8752b2ccf · outbound

This paper cites C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.017306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.017306Z digest=sha256:19202b794379eca2d349ec4f6fbea550c090c648f3937eff5584416ba8d437a6

Observation af5f835a-0a5b-41ae-ad1d-c0f68a006d57 · outbound

This paper cites S., Kim, J., and Yoo, C.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Kim, J., and Yoo, C

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.080901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.080901Z digest=sha256:0c5cbeec9e3d132c2bc7195b03144331dc711ceda34a8047288ee081f53a48c6

Observation 77f05230-c189-4624-9aff-b930a294cd2c · outbound

This paper cites Tpc: Test-time procrustes calibration for diffusion-based human image animation.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Tpc: Test-time procrustes calibration for diffusion-based human image animation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.318669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:45.170101Z digest=sha256:40df0716f472b2e140f6f78b3bc5070c43c1fae1354f57505435ba41787a11ad

Observation bba0a0e7-f868-48a1-bac1-32d45411538f · outbound

This paper cites RRHF : Rank responses to align language models with human feedback.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization RRHF : Rank responses to align language models with human feedback

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.199167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:45.244056Z digest=sha256:055658ebb8c2beb5c57835ebdccf24c32b95181ff822204ccc284f9d06d0780d

Observation 1eeb1d21-fd0d-487f-a370-57350e86ea35 · outbound

This paper cites Token-level direct preference optimization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Token-level direct preference optimization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.081029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:45.277916Z digest=sha256:0ecf53b3f46fa522c87f0ab6a1a07fb6efc60ca5c2b1d844723ee5323fb95447

Observation 1dc57259-667c-4d60-b5b1-f7b8a9949c6b · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.342641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.342641Z digest=sha256:5bc7b3febde9bc0389c85c83f2ee6f6e1fa0427a15ea38afc37446f5cf94764f

Observation 9e796a6b-07b7-440a-873a-5c8207e6972e · outbound

This paper cites P., Zhang, H., Gonzalez, J.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization P., Zhang, H., Gonzalez, J

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.422180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.422180Z digest=sha256:340577e4825e6123f5707f5433a1c622a7666913a730739f07d215e000995fef

Observation 2cb49ef8-9c9b-4f85-9bb5-73958cbf26a9 · outbound

This paper cites T-REG: Preference Optimization with Token-Level Reward Regularization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization T-REG: Preference Optimization with Token-Level Reward Regularization

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:14:45.733451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:14:45.496001Z digest=sha256:ec656557128e4188c9c5f25e21773badf282f134de209b9d6377c925ee3f9a55

Observation a0194a21-5383-4b36-b536-83db9bb9b377 · outbound

This paper cites write newline.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization write newline

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.567616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.567616Z digest=sha256:012c028c735bef62a4d93e3f0e9d4490cb287034a13d82914b9653225e3546b9

Pith citing papers

Observation 71ed9b45-e61e-4f44-81eb-d7947419bbf3 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.830504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:4c3948c628d77e613ad8c5f5f13e3e98e1c3382fc99c4f7d0ec984fa673aa7cf

Observation bf7f5771-d720-4a92-87dc-47e4c6c247e1 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:59.503415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:59.503415Z digest=sha256:6345e8ec74ff5485474965911bb40f762ea0e6e9a9a669675b88a91a1cd6fe95