Pith. sign in

Paper Citation Record · LEDGER

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

As of 12 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.14614.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14614 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:39:11.464437Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0689e7eb-20f2-4117-8bdf-3be15eeebe3d · outbound

This paper cites R., Geist, M., and Bachem, O.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization R., Geist, M., and Bachem, O

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.111981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.111981Z digest=sha256:467ab256ceb687d295f45800ca949565d307af8e7c146cc9b33877bd2e687f38

Observation c3e5efa7-533b-43f0-b14f-cbe4f735064e · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.261379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.261379Z digest=sha256:c935317f6732c517ed88f40a65aab8d34eee3d1c6dcf7c21218c7b3d8d0f640f

Observation 703dcbbb-8474-4336-9b52-e30328945863 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.367580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.367580Z digest=sha256:9aeb70bde3e2b94f09bb0c12a500cf20199f7a4c07a4368faf5d683195a098f1

Observation 61b93e40-3d7f-4cdf-9e2a-120b63a30cdc · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reasoning with Exploration: An Entropy Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.528882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.528882Z digest=sha256:d3f42c227e24151a70d94ef824c8c38d4080caaa1d705d8160809ddc47e0cca1

Observation e9dea355-746b-4272-b43b-a642bb4cdfdb · outbound

This paper cites K., Chen, G., Xu, W., Luu, A.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization K., Chen, G., Xu, W., Luu, A

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.651251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.651251Z digest=sha256:03bcdc07f6107e9fdeb9645bfe7d69324028335ecdcd9c475167295146d29ed4

Observation b230e182-72c6-4661-8c47-925519a34555 · outbound

This paper cites Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.050935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.050935Z digest=sha256:526ff6f385f850eff028192135eb230daa87b6959076a99d59d4745cc2b447e0

Observation 20fe05de-a943-431b-a89d-ac5d7396223a · outbound

This paper cites Deepseek-r1: incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081): 633–638, September 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Deepseek-r1: incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081): 633–638, September 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.256223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.256223Z digest=sha256:04aec82b72ab9cff6ebdee0b8cb8838a295c1399869c08210cf193995e1883ba

Observation f0e898bf-91bf-4164-b613-9c1939a72c09 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Measuring mathematical problem solving with the MATH dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.800750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.800750Z digest=sha256:b1c09887364b1ae7e922f713aa03fca2cb6ee0bdcefc0391d8260b9cf93ff86f

Observation 65be8710-17e3-4c27-89cc-916385f00bef · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reinforcement Learning via Self-Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.961710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.961710Z digest=sha256:0573767333e4fe5869c966ac5f841fbc8eed7e4ae7d0902bc768c48348b5c4c1

Observation fed4f4cb-834b-4735-97b1-c9820b7609d1 · outbound

This paper cites F., and Joty, S.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization F., and Joty, S

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.080907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.080907Z digest=sha256:e0ec555d0ea5ccfebc7661755bdeaadbd7c03d488d5227fd8f34ddee4ef36563

Observation fdd53a9a-c87d-4837-9844-6cd8ab6e4b68 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.211229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.211229Z digest=sha256:69c7c71a6bc7afa451fdef9df217c30a2bfa478db812a8f75931c7599b54f938

Observation fceea892-3f7f-4d84-bd97-d024ff84b8f4 · outbound

This paper cites V ., Jeon, M., Vu, K., Lai, V ., and Yang, E.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization V ., Jeon, M., Vu, K., Lai, V ., and Yang, E

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.336708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.336708Z digest=sha256:2fa9ab6185a1244e9020b3b0e8c0a2bcd6a6df8632fa0c490d548776b79de8d6

Observation 2d47ed9d-be70-458f-ac3e-f1b4e783acb8 · outbound

This paper cites Revisiting LLM Reasoning via Information Bottleneck.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Revisiting LLM Reasoning via Information Bottleneck

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.465608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.465608Z digest=sha256:cd9edf417876991f91a2daaff18c63952b07c6bab51d43f05585e45ead003070

Observation a2ef72b9-6fad-4cf5-bc2f-4e97442b67c9 · outbound

This paper cites RED: Unleashing token-level rewards from holistic feedback via reward redistribution.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization RED: Unleashing token-level rewards from holistic feedback via reward redistribution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.597752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.597752Z digest=sha256:9e7e85e7a7e77fb3c2246e60bd9d74b6ab34dad81955b07eb6aaf0dd02508b3d

Observation 80a54d28-261d-4a24-9277-f7be0d804108 · outbound

This paper cites Can we further elicit reasoning in LLMs? critic-guided planning with retrieval-augmentation for solving challenging tasks.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Can we further elicit reasoning in LLMs? critic-guided planning with retrieval-augmentation for solving challenging tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.743422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.743422Z digest=sha256:af4d660dcad669d5618626e6fc138b2d2f61097f457eb40f719c6d0219e69afc

Observation 71afbd51-17a8-4d77-a281-67b2a3efd58b · outbound

This paper cites Let’s verify step by step.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Let’s verify step by step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.881210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.881210Z digest=sha256:7ede8b02e9a86a39dfac34bb20ee1a03456f7b3a2139d04a392fafe0f83e589b

Observation 29d4ae00-9413-4c24-a698-8104a05dbb6b · outbound

This paper cites Ravr: Reference-answer-guided variational reasoning for large language models.arXiv preprint arXiv:2510.25206, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Ravr: Reference-answer-guided variational reasoning for large language models.arXiv preprint arXiv:2510.25206, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.013709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.013709Z digest=sha256:4eafdf6817db503754dc3fe6bd90424ac7de1c73584e4e33588c7ab0b259b064

Observation 4aadada7-4f06-40c2-ae7f-1c52211b643a · outbound

This paper cites and Lab, T.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization and Lab, T

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.145093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.145093Z digest=sha256:523cb07682fb55334adca24a72006d2341b2aa1e672d479e08d4ff6afb5d43e5

Observation 25553960-5153-441d-a7f7-d972e5747e04 · outbound

This paper cites American mathematics competitions (AMC), 2023.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization American mathematics competitions (AMC), 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.341996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.341996Z digest=sha256:92d060f33143037bbdbd68f542cf8bf3d3816f04a3c6a4752ead60c7d7268887

Observation 378c8d48-05fa-42ab-89c3-d7fbe08828c7 · outbound

This paper cites American invitational mathematics examination (AIME),.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization American invitational mathematics examination (AIME),

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.502677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.502677Z digest=sha256:8f62662f8864117e0f761984964548287ad969d2d95783205a2a4e055bfeb0a2

Observation da0d60d1-89d8-4da0-8fe0-adffee18193d · outbound

This paper cites L., Stickland, A.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization L., Stickland, A

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.860327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.860327Z digest=sha256:330b3c1b668471cac953157e4ee062963ae1cefc05e7c1b4bea807aac7002e67

Observation 0cf673c1-2b74-44d7-a479-ee9a0a320d1f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.049351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.049351Z digest=sha256:08c30c18ac180ba0aa58866acb6389cfc6a27365012d4966ffa89d67cddf2535

Observation 40e7aa23-9973-41a1-bbe8-68ba3ebc29eb · outbound

This paper cites Accessed: 2025-12-23.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Accessed: 2025-12-23

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:06.661112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:06.661112Z digest=sha256:02f8d8602f1a92d181890559269b789324731879f131722ea5bd6c54f31f8614

Observation 61dfcddc-5072-4349-ab74-4c0092e78eca · outbound

This paper cites On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.405827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.405827Z digest=sha256:b48673ad380350127ee162c1380b7ffb6f5651cea8e315f5de7c915a3119900a

Observation 1f33537c-e256-496a-9950-69b95093438c · outbound

This paper cites Self-Distillation Enables Continual Learning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Self-Distillation Enables Continual Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.574318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.574318Z digest=sha256:a5821fac20f12c0b064c6c788ee6fd18d7c63935a376adf90d667fc778a62a0a

Observation e94f6c29-ce3c-4863-a626-f99a505f94d6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.238935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.238935Z digest=sha256:b423c6c131bf87caa11d7ef3d8210326c1e52f8ca99cd337fd24c92d2166e7cf

Observation b1bec4d7-fd43-4282-aa3e-8760191cd133 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Kimi K2: Open Agentic Intelligence

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.864772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.864772Z digest=sha256:dcf41e6d7b06d8f1a98abcd615a9da9abf7103e8e116dc6128bda922a288597f

Observation 87738a82-72e2-41dc-bae3-09dd7cd67609 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.003116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.003116Z digest=sha256:6687930cad2af7c3ba117a3d1440aa9603eff0d75ab806d5933d1ce104b90b1c

Observation ce0c3b6b-14d8-4b8e-9bbf-26743f2329ba · outbound

This paper cites Espo: Entropy importance sampling policy optimization.arXiv preprint arXiv:2512.00499, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Espo: Entropy importance sampling policy optimization.arXiv preprint arXiv:2512.00499, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:07.745985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:07.745985Z digest=sha256:108bdae9e66ed3231b54a34f4a7d6d66ab7229e64014af0e7bb03bc5d0ae8b5f

Observation 9f56af7d-adc5-4061-ac52-35a00dfab8da · outbound

This paper cites SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.323586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.323586Z digest=sha256:fcd5772078ad7a14dbbf06f14d1336b6f5bc6ec3052869b422979f08175efa6f

Observation ff2c5ad0-5fd8-464e-8202-baa1d11ad2fd · outbound

This paper cites Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.449103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.449103Z digest=sha256:c354cb85d7fa8c7be9b5332beb7227f1b786f8c5bb0be91b75776eeaf154c120

Observation 73026038-d1e2-4d6c-926f-ef4f97237b75 · outbound

This paper cites and Karkhanis, D.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization and Karkhanis, D

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.192396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.192396Z digest=sha256:451b38fb330c3c2f53c13fc9dcbe9e1f6d95cf91807ded56c34427276c558ea0

Observation c18f4ab1-a996-42ca-be58-bfbab0daff46 · outbound

This paper cites Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.714342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.714342Z digest=sha256:cb234f671db75b810a4e1b9474cab4bee0bd720b29d935b5abafa63d416b95b2

Observation 83fa065d-7d82-4a64-9b34-c917ca434a1c · outbound

This paper cites MMLU-pro: A more robust and challenging multi-task language understanding benchmark.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization MMLU-pro: A more robust and challenging multi-task language understanding benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.906301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.906301Z digest=sha256:0b084fcf41f1827a198286a6a189f60525f775a24329fb8553ed3c944fd47e7e

Observation 8f2f0074-1d9b-49d1-84ec-92c6fd15cd36 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.605142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.605142Z digest=sha256:049fd90c54469992532c18861965a3387c173b9e64938e1079447ebe998adf9a

Observation ddb3deaa-f872-47f7-b5f4-6ecc6f2c91f6 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-02T01:39:09.162940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.162940Z digest=sha256:40c3b5ae7f68d80114e54844e6706d30163a0a98541d4a58e7d1082578e494fb

Observation 52a76309-0236-4747-90da-a17746c21878 · outbound

This paper cites Quantile advantage estimation for entropy-safe reasoning.arXiv preprint arXiv:2509.22611, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Quantile advantage estimation for entropy-safe reasoning.arXiv preprint arXiv:2509.22611, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.291533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.291533Z digest=sha256:0fb65bff85be81b4f23702c37bcea82058a7c8fade6cdd3c79a19131d528fa28

Observation 43079864-152a-4fda-b2cd-4706f00e4322 · outbound

This paper cites OpenClaw-RL: Train Any Agent Simply by Talking.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization OpenClaw-RL: Train Any Agent Simply by Talking

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.040248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.040248Z digest=sha256:16bdab9aaf2d52e50b983ac32ccc80a02df754d626bce550ee2e745534e9c9e4

Observation cf97f670-e25a-4dfc-bc18-fbb7f0b853f9 · outbound

This paper cites Reasons to reject? aligning language models with judgments.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reasons to reject? aligning language models with judgments

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.686011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.686011Z digest=sha256:1aed9378e3ef97249825082d3b74cae268ae817fee65b5044ed97e2c4818742b

Observation 804dc322-12c0-41bd-b001-027a18c64b29 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.838682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.838682Z digest=sha256:5370a88ff8c9e711ad2ede18afde80e1a1bd7c7b08e9e86c064f3caa9db71f0d

Observation 5017778e-003b-4130-a602-849d58ae4ddc · outbound

This paper cites A., Osten- dorf, M., and Hajishirzi, H.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization A., Osten- dorf, M., and Hajishirzi, H

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.482306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.482306Z digest=sha256:24524640f3a698e65d85640573ea53987d8cca333174e3015f23f31edd42e3cc

Observation 30e4be79-6d98-45a4-b9e9-669e3bc40d89 · outbound

This paper cites Self-distillation bridges distribution gap in language model fine-tuning.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Self-distillation bridges distribution gap in language model fine-tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.134964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.134964Z digest=sha256:000a172b5e5fda8f1e7d9ab37a6fb8c1d0f9b2c76ce454153d78dc07be8b7ddd

Observation d0a49c8e-a0dc-4b59-b561-71e72d61c698 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Tree of thoughts: Deliberate problem solving with large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.331443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.331443Z digest=sha256:0cb93bc92be61f076d81517e18042cbbc6c7df341a6e6ed0938298c5448f0493

Observation 6c38c0ee-0f90-401d-af90-170dc0ea9620 · outbound

This paper cites Qwen3 Technical Report.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Qwen3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.987173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.987173Z digest=sha256:92022762890128e5d701f54b8d8266ecbb37da9fb6c99ff4dd228796aea6c721

Observation bbdfea30-7b90-42f2-9959-1aa0c07b2398 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.461306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.461306Z digest=sha256:3eb3a5828f617c18aad1a14c2afb65db85063c87258aa338de5d9af51f6a909c

Observation c46e2e8f-02b8-4482-9d30-57bcf13ee76d · outbound

This paper cites SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.555531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.555531Z digest=sha256:94e1f03f9a1a85ed60572a3945f3c34d01b42ed07cff47529de3aea40ab42672

Observation 17252ee0-4e5a-4a37-bc01-6e96b3b77dce · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization DAPO: An open-source LLM reinforcement learning system at scale

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.410737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.410737Z digest=sha256:57441fee5e550a324384c7becc6215890eabb1cfb2977108a4f64776d5b4a0d9

Observation de45ab0f-6525-4a0b-8a4a-4ac946e6c732 · outbound

This paper cites R., Kailkhura, B., Lai, F., Zhao, J., and Chen, B.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization R., Kailkhura, B., Lai, F., Zhao, J., and Chen, B

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.749976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.749976Z digest=sha256:06434d06d0e196d029e29cb784256267ab1b343d1f1b7fc98cdcb1f3659d56f3

Observation 8331d4b8-dc11-4eda-b24c-c560ab2b1b34 · outbound

This paper cites First Return, Entropy-Eliciting Explore.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization First Return, Entropy-Eliciting Explore

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.853118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.853118Z digest=sha256:d3e6b828b3c8c866e65d8de4f204e60d5b84eaf7a36673da46246ccf6a6b0931

Observation ee0017c6-1e68-410b-b189-a4cd75cdf2f4 · outbound

This paper cites Group Sequence Policy Optimization.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Group Sequence Policy Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.657660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.657660Z digest=sha256:907ce013f17e24567aaae7f3108e87db8d305405e4de114d7b767826d2551b3e

Observation 2d65af65-e512-42f9-b615-bd93fd42e247 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.953149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.953149Z digest=sha256:b5620c41e02366ed82431aa1f30f59880dd2788f879a2d7ae2eab824a281ccec

Observation 4291b401-e973-4b24-aaa0-f65e5ad55dce · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:11.054736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:11.054736Z digest=sha256:75d271018cdf1bad37821158757989f4a9ca62c451e2d798b9ce7a59467f4e99

Observation b3aae2d3-5a8a-4f6a-97a7-4faf062ba1d2 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:11.181475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:11.181475Z digest=sha256:d546a696a3fdac830005eee4ae5ad9a6c5418ccaa9cbe99850f65272c70b3b69

Observation ef22e5eb-c616-46c7-90a0-0193d8386361 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:11.357786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:11.357786Z digest=sha256:15613028f42ce8f7876922f7365e53fb622844da682d3479b3afbe93a6c64949

Observation f6baacee-6734-4eb0-a29b-3a115665d2f3 · outbound

This paper cites an unresolved cited work.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work

Reference 58

Resolution
malformed identifier
no resolver link, observed 2026-08-02T01:39:11.464437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:11.464437Z digest=sha256:c1e4a0a82eb51cee00f3237608bc1c516735c1dc72a16c1be3da6faed6bbdd5a

Observation 064e667f-ead7-4b58-9de7-827eee6bce35 · outbound

This paper cites URL https://aclanthology.org/2024.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization URL https://aclanthology.org/2024

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:03.851537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:03.851537Z digest=sha256:7d4fcfc0fa227ec15b04109dfc9931dea542305f891b69b2fd38bec61c6192d8

Observation 4e3e6b8c-43f5-4ad2-a9d5-4ff845c7fd23 · outbound

This paper cites Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:04.601106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:04.601106Z digest=sha256:ce09de3f1e22dcbdebee1fbcb5036b4fc5cddadb09b5d692a97e02839d761b74

Pith citing papers

No inbound Pith citation observations are available.