Pith. sign in

Paper Citation Record · LEDGER

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

As of 13 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2505.23540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23540 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:54.646894Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T02:10:40.020460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T03:55:53.842268Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d0201b1-c789-4d40-a1b8-4531e2a509e7 · outbound

This paper cites URL: " 'urlintro :=.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.753612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:46.753612Z digest=sha256:5efb412950b50dfde7167f900a3357354b2b2ea7faaab760d58edb0fee6a5968

Observation b9fa71cd-82ec-4690-a49a-03bb98b80ddd · outbound

This paper cites write newline.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.884044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:46.884044Z digest=sha256:4f1950d369d8668cf9aa0558c69bd7894a0d7168c37fa41c5b50df1526350c6f

Observation 2cc13414-55b7-4a9c-abf4-e85bea0b9a3f · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.051024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.051024Z digest=sha256:f810f27ffe86d1710235e1a96646c20a81ab00ab74b4a2ab503fca56e32641ed

Observation f36448ab-9dbb-470e-bfb2-1ce6b9b67b6e · outbound

This paper cites PaLM 2 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning PaLM 2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.263325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.263325Z digest=sha256:bb6398e96bba46561709708fad57819cb35a75d8b6067f3c35217a7d2404d8d2

Observation ae85c740-cb08-4a98-8fe2-77b25cb86a95 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.435160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.435160Z digest=sha256:a8cf063ac532faa6d25de46d86a095a7094a2b9906570ba5dbc0dba59d93d6d1

Observation 0a5a5b99-02cb-4289-8626-cc777c298547 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.525774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.525774Z digest=sha256:fde5ad773a1b52c6fe28436b04eef70d4fc51dd7b2d39f110c0769b100e70ce3

Observation c877b047-7045-4853-8417-22c2d296aa21 · outbound

This paper cites Qwen Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.699588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.699588Z digest=sha256:9cc2c9da145e7db14b9792bdd3399f872c8626e63c027c67cff9575b591cd814

Observation f0670dc0-94b8-41d2-b2bb-c56c94912132 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Constitutional AI: Harmlessness from AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.872992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.872992Z digest=sha256:9c3ee59a9aa1cf7f2c310fb5519501b092871879260a656c1c1de35a8a68572c

Observation 4e73070a-331a-475f-9fac-195c6406a8e9 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.005985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.005985Z digest=sha256:e69bf790d320954e3d45f51260cf56b30a867debee7fc0f1bc60de1e9b70ff3c

Observation 789ab79c-abe2-4128-a481-15da00115460 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.128007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.128007Z digest=sha256:65b963592c216bdf888a9a5921e134b5ba5e3e1645b445d53d7e8a9c663d3a02

Observation 04de3fd1-ae6a-4501-9a81-4e0544bd0d13 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.284922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.284922Z digest=sha256:fde15254757623af078d97641ec44bd281e7c88cd1666d5f8a1ad9f3ff05eb69

Observation fe43d705-b098-4d91-80cc-20a738956ba7 · outbound

This paper cites The Llama 3 Herd of Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.497500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.497500Z digest=sha256:84452888b9f3d23dcf85e5046b274ba6d68830faba830420972d04520a6c7748

Observation 5d37c2cb-3193-40da-84b9-393519fa4a1a · outbound

This paper cites u rnkranz and Eyke H \.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning u rnkranz and Eyke H \

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:56.641896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T12:48:48.600730Z digest=sha256:2f121ec497927a817a70f2d4a723f83571779613b5ced181f1b2adab8a38e8fe

Observation 9ed398c0-88ae-4c62-87ff-d30bc9adbbb4 · outbound

This paper cites Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.773375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.773375Z digest=sha256:def3d8f6e9e50f32105377e6c150d8cdcf181cb154a787e87c4256ed601c2457

Observation 807d9afd-850a-474e-a795-f9013df45462 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.941524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.941524Z digest=sha256:67486a9cc9b20db1857110e3373009368b221eb2e394965adce9637812eb58d9

Observation 1d2b6096-c2c2-44ea-bb5f-f08a890ccb5d · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:56.412047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T12:48:49.082974Z digest=sha256:cd4c89718cbae782c388bddf3789cfaa9d6b4cdb512dbeb26fd254a4228ad8db

Observation c0be4535-8a3f-4e1d-ba62-e0be2797985c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.267026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.267026Z digest=sha256:42898c9bc12e1549d6fbac65d24daf000587d8d1e264c21c15687c2c5b0fa2b8

Observation 971a7edd-e295-4cbe-977c-2b9e0b5e304a · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:56.170201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T12:48:49.366558Z digest=sha256:e8017143bad11d3322dd0beba5640ce64178cf65b039626f6dcbb65715724c1c

Observation 9a896b54-0054-41d1-b37b-f21852b92d9d · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning The Curious Case of Neural Text Degeneration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.542632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.542632Z digest=sha256:5d80272e69f0579c360a44431cd40bc900b3481514b803aef52a40a435aad8c4

Observation aa60184e-e272-4d86-8765-02f23d4f1a18 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.764910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.764910Z digest=sha256:92110d72ef11122ea4f04613a878272e856b21247fca5a5c99032e6c26746a5e

Observation 8589f443-2d0a-49af-a999-e6da8d773234 · outbound

This paper cites Mistral 7B.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.921998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.921998Z digest=sha256:85ceb09f733c1210c011baec6edca03cc834c887c27e28bdcceca34b1d5018a6

Observation e39e65c4-b8a8-4410-93bd-1a9090af818b · outbound

This paper cites Mistral 7B.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mistral 7B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.073075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.073075Z digest=sha256:81c2ec9410d337ff2e898a47efdfbec9e703fd95c3efbe175fafff913b2846aa

Observation a93cddcb-a3e8-4f57-ab22-35ccfae51949 · outbound

This paper cites Mixtral of Experts.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mixtral of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.236031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.236031Z digest=sha256:9ff7cc0c3281037e8287681d4b2ba5016f06943617edead9757e1b16ecf58c55

Observation 5d64946e-9d68-4229-a38a-d70737774dfe · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.415748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.415748Z digest=sha256:1324a27acde94456c39ea2c16956a41b58f7dc1ae6e0efceade56fdcc322792b

Observation c18b59a3-11d1-4f60-a310-2f278c5be249 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.896260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T12:48:50.526871Z digest=sha256:cbc19720793caedf298397acef3e9cca4dcb89b016e18af91622417f32b4ae64

Observation 25372d40-7a3c-47c2-9c23-05b9270628b4 · outbound

This paper cites Let's Verify Step by Step.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Let's Verify Step by Step

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.661025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.661025Z digest=sha256:d7311df250503a84fb415da3aaa84619627850cfb706491d7f084a5be347ce58

Observation 4088bc24-c3d9-4787-8e90-48c51ab95d93 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Rho-1: Not All Tokens Are What You Need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.753591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.753591Z digest=sha256:aea58a44c9a8bdee4c14fa0bb773351f124cc933e05344ee09484055408146b9

Observation 4dc331b4-4341-4494-8e12-08eb8254157a · outbound

This paper cites Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.846271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.846271Z digest=sha256:9bd1f340716c19dc86cd3bad420c48b5e4abce759635949a7bb163c3b140b637

Observation 68230bdb-df1b-4f35-9f3e-b72edb08be27 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.654117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T12:48:50.956262Z digest=sha256:10096b0954ae227c8a7047ed18442304a067a9881f983e7db100af54a3e1de61

Observation 3dc8d264-c4ac-4339-a805-f8f2b22d8a18 · outbound

This paper cites Inverse Scaling: When Bigger Isn't Better.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Inverse Scaling: When Bigger Isn't Better

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.113697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.113697Z digest=sha256:6a50ceeccbe54a36305bff403b259f137b3ca07300b7d634c18632cacafd3037

Observation 292d24a2-6db8-4036-bb00-29d468322805 · outbound

This paper cites Large Language Models: A Survey.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Large Language Models: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.344663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.344663Z digest=sha256:e48b04b36d5c3431caa3382b9de0578e637f8635d3d098acb6478838539d67d0

Observation 84a2e209-521b-4d8d-8071-a0a6fea3e594 · outbound

This paper cites GPT-4 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.544455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.544455Z digest=sha256:cc157cb62e02989406866ed6a311186d794b86cfdf51bda5eac290f4e74a3388

Observation f24bd252-aa67-4027-9ffe-31d4bcb8ed9d · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.715839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.715839Z digest=sha256:7d76ddfec355d6b2f6d6651442e7ead31b24b644064767272cfc53c418ff22ec

Observation 6b8c74b3-c7d2-4fd9-9fa0-e88b8aaf2d78 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Iterative Reasoning Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.881145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.881145Z digest=sha256:cf3e3ac2d0b89e7160ff7da83c21542ce7c3336de8c2aa98eb3f225ea3f68f99

Observation 7c5d64e3-bd74-4310-87b3-46498df3ea7f · outbound

This paper cites Self-Consistency Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.028303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.028303Z digest=sha256:b226993e33efb7e562c2db0d748ccf555a83e3fdb7ebab02f282aeb31f728af5

Observation 63350242-06bc-4505-9b17-e3593cbb3996 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.236313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.236313Z digest=sha256:70d1b8cd30aa69f32c621cc1f8492936845f71f3e681f57f015a18335023677e

Observation 5d901273-a9c0-4812-90d3-7e8131a2b228 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.442199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.442199Z digest=sha256:8eb206a7704824e0bc841a539e30ba0acb9860c05f8f3d001321689afc62ed11

Observation 2a5a65b9-8bbd-4185-9e6d-cf89e8a9d320 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.594267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.594267Z digest=sha256:d890df6895a7a83a673c3b91b3ce6e6d5285b1878f51810a8f6d977f819a643e

Observation 5628600a-17b9-4f11-9666-e7cf21c8ed3f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.739102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.739102Z digest=sha256:81fec28ca968535e9822c3f8afae285d6adf78712a464aa49ae6e6f62afdca13

Observation a53b2819-2362-4d2d-917a-ff8642221b71 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.883958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.883958Z digest=sha256:9154f8bfe1174d39039dc5b655ce52f3788b756bae776fef2f7e4f0734e85d21

Observation e459c472-4f91-4d75-b669-3e9993087f5c · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.070330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.070330Z digest=sha256:077abb5b31cb5fad0292b2048a9b8b31c642c9e7ed80f1131356aa8a04ca8900

Observation a118128f-1d83-4816-a98b-b1ee6b2c5513 · outbound

This paper cites LRHP: Learning Representations for Human Preferences via Preference Pairs.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning LRHP: Learning Representations for Human Preferences via Preference Pairs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.221253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.221253Z digest=sha256:79f057ed3bca722ff974bb56114eb141b2dde911ecf591119da5c5167fa91ed7

Observation ccea8820-6185-4dab-846f-fe5503aacfb9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.362214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.362214Z digest=sha256:99387be54a23247570caa8e08339f9795e25f8eb3fc435939685ea5e4a13f317

Observation 7009d628-dde1-4970-b138-c3801022ed5d · outbound

This paper cites Consistency of a Recurrent Language Model With Respect to Incomplete Decoding.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Consistency of a Recurrent Language Model With Respect to Incomplete Decoding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:55.005187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T12:48:53.508094Z digest=sha256:00ea87144eef3b6829ba2cd232535e6ccc343e59c7ddf9e8ca73ee20cc7321f9

Observation 6261dd51-e4f9-4d11-89b5-626a1f7d92d9 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.456689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T12:48:53.657068Z digest=sha256:92c5c27d3e68f64b6780db6e6dbfb389848555d5f44f4734ed53c7cb86428ed6

Observation fcbbe315-6b07-4e8d-be48-ea0ba4728290 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.850737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.850737Z digest=sha256:42766f30dc71b5acbbcc6b3d4a5d29ead8be480daa1021ffe7f01f18a40d8183

Observation 0cc5b074-990d-47eb-8731-df040fe7866f · outbound

This paper cites Qwen2.5 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.971380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.971380Z digest=sha256:c007eb59dc8a46b46649ff68de9636e86f85aca22fc9bb0806a8e7f95501ae33

Observation 12fac24c-5d97-4132-8ce7-02079af468ed · outbound

This paper cites Qwen2.5 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.069667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.069667Z digest=sha256:c7be23c5f62facad8f9867d0400de4124aaec6781e71a48766eea83b010e944c

Observation 0c0ef65f-e24d-4869-a4bc-854791757086 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.203982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.203982Z digest=sha256:0ddb0d298e0a4920775d81e95d9e3ef4e2edc646e28befd081df146a2089466f

Observation 37060bab-bcc1-490a-83ca-f6405add8166 · outbound

This paper cites RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.380705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.380705Z digest=sha256:febb0e50e23094f68dd28cd426be12c14f478481f4f73f8dabc979e27371485e

Observation ca19ee2c-5fa0-44d7-a150-b0197044a3a3 · outbound

This paper cites Self-Rewarding Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Rewarding Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.498345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.498345Z digest=sha256:29adfc31c7e8c28e2cf05724a0bbc302faceb1b52d12acab3c628e71a47ded8a

Observation 8cc37db7-4a5c-40ce-99ae-84eaceba409b · outbound

This paper cites Token-level Direct Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Token-level Direct Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.646894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.646894Z digest=sha256:b6030d5c27c311f215104c7cd370f382543b861e2dfb99a7ce3e93480d3cf4f8

Pith citing papers

Observation 704f80b4-7721-4f40-80da-ebb3b85067ed · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.845454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:7f978f31aedf8743773e0cfaef4490ecd197e20f0b9a7cafcafc853c57659031