Pith. sign in

Paper Citation Record · LEDGER

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 14 inbound Pith citation observations for arXiv:2509.09265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09265 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:29:01.480681Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:48:09.620197Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.648363Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8aed9b11-2178-4d64-b947-6ad892f692a1 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.091838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.091838Z digest=sha256:16a36a02ef73f60bb02f22f9308a4d25c903c9f244f0995bf894a58b3ca1d76a

Observation e6144339-f32c-440d-a59d-ee4fd68eaee3 · outbound

This paper cites Open Deep Search: Democratizing Search with Open-source Reasoning Agents.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Open Deep Search: Democratizing Search with Open-source Reasoning Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.149783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.149783Z digest=sha256:15914808f97695d732d2fdbe04e87ca3c24c0316b1f73760c3178b93183d48b5

Observation 69199cbf-c8ef-4641-92ed-6d7e0819c1b2 · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unifying count-based exploration and intrinsic motivation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.219728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.219728Z digest=sha256:56c68942eda7a92ebbf442d290c66b0ef324997c5c6e44d63377d8793c8ef41a

Observation 2e325a43-c496-45a7-97ee-128721ef66f1 · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.258016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.258016Z digest=sha256:f875ae9b24aa8749959ee97d3e796c934ef4b6c20eccfde5f895a0c2c459ca81

Observation a05d97b5-0110-4c94-ad1b-28ebe8aa8c2d · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Reasoning with Exploration: An Entropy Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.355996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.355996Z digest=sha256:933ded91edb6c8d7e930b906c5645d2164a9bff372ab860cf53332aaefabe286

Observation de68e66a-70e6-4029-bada-9fd30d5028ef · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Mind2web: Towards a generalist agent for the web

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.497613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.497613Z digest=sha256:46aa8fd464fc9aca5c4ce4c1fa0424c8e4b260c1a3bfa74e065d57259e7d6f83

Observation 90f14950-d6d6-49b2-a74a-88977f7ab7ac · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.630716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.630716Z digest=sha256:941978ae2fbdf4593050ae3bc6b2498adcdf336f94fa033014515156e3c5dec1

Observation df389e07-afe4-4d95-b7c9-51e573fb0dbc · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.713427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.713427Z digest=sha256:dec32f3a71a95b6a8fd322b1559f7fc80ff1fea1c8398496ee276311ae62e1b9

Observation 596ef3ac-f757-494c-88ad-d727917f3bec · outbound

This paper cites One-shot Entropy Minimization.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents One-shot Entropy Minimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.841114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.841114Z digest=sha256:fb5c689d5fd71dd013e7f6ab14a940308175be893c25fada89715e78c18697de

Observation 5c180ad2-db6c-4ea3-aa5a-3719bcc5a038 · outbound

This paper cites PaSa: An LLM Agent for Comprehensive Academic Paper Search.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents PaSa: An LLM Agent for Comprehensive Academic Paper Search

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.997117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.997117Z digest=sha256:2365cf1e4cf7f5245f782cc097b152cae6c980b5c6189e82073fd41c8143254c

Observation 8a1376b0-98e0-453f-9a7b-6f45d307f65a · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.112469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.112469Z digest=sha256:026bf28f71308eb0cb7c82461ba9fccb2c6d4e52063cd49a0cb33dc81a5b7522

Observation a98c37c0-4a0b-49d9-aec4-9c9c63dc83ec · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.248343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.248343Z digest=sha256:1743153fc88c91584687c3e648517c88325c91347f3954d9947892c98fa81186

Observation 7a2f0c0c-7b53-403d-a0a3-9166689d0131 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.347328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.347328Z digest=sha256:4e35f1aca5d51cb71dd1af63f58ddb6078d0ede737558f1d8d92c609f2b99014

Observation 38dc535f-1696-4f01-833e-38b59d342411 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.447645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.447645Z digest=sha256:e064565dbd744bd7d5485ebc6b159540f075ce708b37dfc4e8d8dbb182a2ffd7

Observation 2fd13494-fcc6-4011-89cc-4a00c826e275 · outbound

This paper cites Vineppo: Accurate credit assignment in rl for llm mathematical reasoning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Vineppo: Accurate credit assignment in rl for llm mathematical reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.566435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.566435Z digest=sha256:2bae15db06f56b2b2d2aa3478bfdce2c363b9d433ba7098ec8e78384eb75e96c

Observation 30f6dae9-c4ec-4c99-b9ca-9a868bbf0e88 · outbound

This paper cites Empowerment: A universal agent-centric measure of control.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Empowerment: A universal agent-centric measure of control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.653021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.653021Z digest=sha256:bd3b773b5e1b5c5fcac7f6e99a3562af6486de2d2c9efad6a0bc9199edc9b1a2

Observation 35a7abdd-7c0c-4d50-a6f7-1171953a310c · outbound

This paper cites Natural questions: a benchmark for question answering research.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Natural questions: a benchmark for question answering research

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.767525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.767525Z digest=sha256:c209648e4edcd0a2b65dd7ee26af693d4a8143f9d6b78564a1a3b7bd613c10c8

Observation f934fa58-2f68-4fe4-967a-97c41b32e0d2 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.876860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.876860Z digest=sha256:2416b76445c4af5f537f16f0924b196d13297b19a64c1730f4315eb16af6e006

Observation 1fb21e88-50dd-49a8-8d0b-b92dd067e492 · outbound

This paper cites Logit Dynamics in Softmax Policy Gradient Methods.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Logit Dynamics in Softmax Policy Gradient Methods

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.006560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.006560Z digest=sha256:2c826325ba48bd61f10738aae388950590d3bc2f426c7b8fa5bc1783797fb5f4

Observation d76f24e3-2abd-46f0-a4c9-b1c56c5755a7 · outbound

This paper cites Let’s verify step by step.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Let’s verify step by step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.159200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.159200Z digest=sha256:3e860ded7d5d24886ac54d0151797ccf4bbcbdd8f4608bfde4831484b961d7fb

Observation 13723c2e-df2d-4094-92f3-2598b454f8e0 · outbound

This paper cites Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.286475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.286475Z digest=sha256:93e9ec9dcdde47fc4f566a4fa8f89e2f40a8a012da757c958fa5a937081538d1

Observation 9f632c4f-3d97-43b8-84de-1aa3165ac244 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Policy invariance under reward transformations: Theory and application to reward shaping

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.398421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.398421Z digest=sha256:af8736a998fa90c9309a9bc7126d3af14d42ceaeb77027896d7596c84664b720

Observation c6dc47bb-497e-43a0-a5c0-29aa152bc314 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Curiosity-driven exploration by self-supervised prediction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.536248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.536248Z digest=sha256:f53d8605ad95377af92d0a67653c67f49fb602014bb5bb6592b57822d9eefeb1

Observation 338e5196-915e-4c4f-a895-fdc764ad3033 · outbound

This paper cites On measures of entropy and information.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents On measures of entropy and information

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.698098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.698098Z digest=sha256:ad7ce8b2f3124d94f459e6d1197fecbeff0b33c001b49029b3823c2bec642659

Observation 496537d7-1db6-4c64-8b0d-8c1941c8e06e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.804689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.804689Z digest=sha256:6d90a7347bbedb4946973b39404394e2f3c09abedbd6783ec2819c33b9760733

Observation c13f59ad-c478-4952-8302-f56a62205fe9 · outbound

This paper cites an unresolved cited work.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.939607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.939607Z digest=sha256:43d48b79ebe5161019477adfbc577e86412eff6a61e98091cf8d15d3a944afb0

Observation 578313fc-82c3-4f14-a902-e6c44baed511 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.039537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.039537Z digest=sha256:16e00970350cbe631c01a2911fcc48ca04ece04b61537bf6f127a6bbbe37cbae

Observation 4cbfc7c3-90c2-4ebe-9eac-f4d5ecd9477b · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents HybridFlow: A Flexible and Efficient RLHF Framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.160640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.160640Z digest=sha256:d55c1e682c39afc71740661a59c2056daf007ca719ff2e94418037c312d1ed32

Observation e2dafed0-13fb-4382-8f5c-c771c102a2c6 · outbound

This paper cites Alfworld: Aligning text and embodied environments for interactive learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Alfworld: Aligning text and embodied environments for interactive learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.252460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.252460Z digest=sha256:e4991d45232ce23e40542ef1ab10d2161d3856826ee72a9a4bc4777188e9a3dc

Observation f6552af1-c61e-4232-b173-973651c25c94 · outbound

This paper cites Entropy-guided sequence weighting for efficient exploration in RL-based LLM fine-tuning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Entropy-guided sequence weighting for efficient exploration in RL-based LLM fine-tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.417616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.417616Z digest=sha256:d57a09277f8ad281bdf1561424845baef4acdb499744d95feb5100c03bae285b

Observation 63c6d6a4-c0da-4806-bee3-5efbee141112 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.532301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.532301Z digest=sha256:622f633ec55b134b3ca6a02bf2ec728ca9915ab757f3f74596bff7fa3c55dcf6

Observation 6e5ac1ad-1a34-4f53-a377-5f6ec8f6946c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Chain-of-thought prompting elicits reasoning in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.677626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.677626Z digest=sha256:852803e3695647523167a65792e34d6b96a5391d1d2450815dd837ba31b8b323

Observation 6515c5bd-58c0-4cf4-b7d8-601ecd8d0062 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.797922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.797922Z digest=sha256:478b9ecbbd20fdfe2ad5b29cb61c0171d2ef645e2c73373b1d7ac3da539d28ce

Observation c23d4d99-89c4-44d1-af99-5536df5706c6 · outbound

This paper cites Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.907127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.907127Z digest=sha256:83615269fa4d161bb35e42bb47bcd02386e7ac8297c93d78269ab01159170130

Observation 9524d2bd-fded-4fd1-aa92-d028f2277706 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents WebWalker: Benchmarking LLMs in Web Traversal

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.015869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.015869Z digest=sha256:237d186a96b4b65d8fd2344fc1ec4456fce07d357366a19cd2b95fad00dd8627

Observation 79d06aa4-e3b2-4652-8bf4-333ada1bd8ef · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.066242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.066242Z digest=sha256:1003fefe9d012ae0fe127540fbb75cb8bf6bf0f1d8b825a7f3ee35944450976f

Observation 6100cb95-1387-4902-ae6c-048d1603b7c8 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.156802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.156802Z digest=sha256:48fafc3d5e6f2475d8a3a48dbac8723d7e4a9391877ce752dcfe83a1bfbc3d75

Observation 7b0c2e6a-f183-4658-8568-718b3f7dc21e · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.216580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.216580Z digest=sha256:d713b72e9dfb511c3b475c53547c62825ce536db1b18ff0e287f0d40a7414f7d

Observation 6f561e9c-1096-4605-be83-fe43692f07a8 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents React: Synergizing reasoning and acting in language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.260395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.260395Z digest=sha256:5f4034cfa10727270b117b98319c6616ea429f784593acd6c5d519b7ad9a1848

Observation b0a364ad-2e6c-41c5-acfa-bb7a56632527 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.312931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.312931Z digest=sha256:2fcfbad10080a02dbae816ad1c343fa2e687f600ee0e127a3507595b645c9d8e

Observation 93c12cfa-257a-402a-8797-7fedceea0e0a · outbound

This paper cites CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.384563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.384563Z digest=sha256:eeeab442c6a2749ced1b6342c989f2cc2d09fc818e68c5f8a6018feb621b50be

Observation 9cadc56b-68ff-4aa3-b874-4412ffc301f0 · outbound

This paper cites Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.494940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.494940Z digest=sha256:38f6d5315ba2c92fa9befd7a97eb7ec776e094ba53c2cb608d162979793234b7

Observation 697a7a5f-2530-4662-8311-315ed08504e2 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.555925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.555925Z digest=sha256:15ad4a26a4d99891d087bcebb68010d9aa195a9bb5439a754238421d7b982a70

Observation 68eb5094-2f5b-4c44-a765-9f568ccab9f4 · outbound

This paper cites EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.634858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.634858Z digest=sha256:6204d22ce622fade813a42fb8cc3969425e7485c55f69befa5b4522272754967

Observation d1da1519-3ae0-4a72-9b73-a586bbd6349b · outbound

This paper cites Siren’s song in the ai ocean: A survey on hallucination in large language models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Siren’s song in the ai ocean: A survey on hallucination in large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.713218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.713218Z digest=sha256:31dc16814632ec128a3f88ce865864f329d750016e30bd2e5861f680b119bf32

Observation 3fcd945e-bbd1-4d25-9ba4-34ab0beecf44 · outbound

This paper cites Learning to Reason without External Rewards.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Learning to Reason without External Rewards

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.812232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.812232Z digest=sha256:03c313af974e2718d3541627b2cb39cba7d30dd852a6e16940363d63ca728275

Observation 8f14c444-b12b-40e3-bcc2-a2bb4ad25b0e · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.879986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.879986Z digest=sha256:0a3d543ce749e2fd7f6625e4c7cec154c39e2a10aa94d589fdcd2aa0203eb436

Observation b62c2e21-5fd7-44e5-b1da-582779056cf2 · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Maximum entropy inverse reinforcement learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.944566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.944566Z digest=sha256:418201a6fea65b4bd700c44b073e93a93cb9126461da2ed985dd24e8137df1d7

Observation 499563ec-ca50-4acf-9bad-8a1cdb473d1c · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents TTRL: Test-Time Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.018560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.018560Z digest=sha256:7f72a9b9b76f8a361af3c553522644e883ec8927aa133a65e8d625ab7509ad58

Observation a1e2fe69-8840-4377-82db-ba5776667af7 · outbound

This paper cites an unresolved cited work.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.084911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.084911Z digest=sha256:a2eaab5e06d34cdb692f56e9d3a6a07c7cc76e60241b0b3b43453f3af96574f5

Observation 13689775-68ca-4d31-bf74-5c608d409620 · outbound

This paper cites stably all-correct.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents stably all-correct

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.211245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.211245Z digest=sha256:a009e2080e8dd062ef32b2a459d730cf3985c0e93a10439196c426222cebcb9c

Observation d4b72e7b-3861-44d5-96cd-595c8510e0f8 · outbound

This paper cites assistant.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents assistant

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.262634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.262634Z digest=sha256:e4c3669e487c1483924c2ea53fb9319950e92deeeb03d9d937b172534dbe355c

Observation 544b829d-5b1a-42a3-99a3-714b0b53bbb8 · outbound

This paper cites an unresolved cited work.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.353844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.353844Z digest=sha256:7161cf173b05c56a9dc48aaac58619806720162a10fb21ed4e1d454dcccfd350

Observation 3a32333b-95bc-4e91-b1fd-9a0f0b8f2796 · outbound

This paper cites For each step, the advantage is scaled by g(Ht) and augmented by the future clarity bonus ζ·g ′(Ht+1), yielding the modulated advantageA mod as defined in our main formula (Eq.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents For each step, the advantage is scaled by g(Ht) and augmented by the future clarity bonus ζ·g ′(Ht+1), yielding the modulated advantageA mod as defined in our main formula (Eq

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.414578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.414578Z digest=sha256:8e6345c944bb33bffa84510d234df90a8139e4373f741597819cf6c57cf9ceaa

Observation 9322fcd5-3d85-4ad9-a19e-fc4273418d69 · outbound

This paper cites assistant.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents assistant

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.480681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.480681Z digest=sha256:b26697e4b55b5a7bb4e8759088e76262f45549620a759350e1bbcef69b08306f

Pith citing papers

Observation f75f5912-6283-4211-b65b-87c1c27300ea · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:09.620197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:09.620197Z digest=sha256:5f6d8df6064321a2c44d773903edb2789d0ded7de6d56c1bf00453b05799a275

Observation 35632d2d-07c4-414f-ab1f-debbc018ba5b · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.221330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:5fe3b2da9651f0d829c7e6dd9d1ebccebf4a85d8cd53382503524c8ce7125d6e

Observation 19f3b51f-5f6d-4cc8-9d2b-7ab238992ac9 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:45.936142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:c0ed04f3fe3441562082ce63c8d8491c4ed24449cde48845b14fd48c1463c58c

Observation 502ff5c2-6737-426d-b756-4871a14cb9a7 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.612400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:1507f6d4b27bcdc55df9683c36c3c5e8987a298aff2378a47f9a9f888d44ffd9

Observation e34849f4-fe1f-4457-88a8-dfb5b5a9d82b · inbound

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning cites this paper.

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:17.685450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:56:27.170349Z digest=sha256:8ecf1bdcb1b621a29b51f262bc0b82fb4ce33688244b72c5bc04d384a6c6474d

Observation e91e85ca-8c04-49b9-83dc-32d945d0a665 · inbound

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning cites this paper.

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:59.145694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T00:58:25.685484Z digest=sha256:e47fd4037418870739007731c84aa7c070c8b066d83f3d27ec3f679237d2c87d

Observation a537d1f3-4839-4130-a9e5-90f74f09a3fd · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.203001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:1d4ea5a381caeaf110ea8bdee697b242bf7c267596f4a9e21a3da19ae27853e2

Observation b267031b-ef27-4861-8b4f-2e522b92afc7 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:08.516567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T10:23:52.522238Z digest=sha256:a7675cd5d5937463e73bcd7892c7eba7ef6e2604c4281ecdac54d5647d39b4e4

Observation 88280954-5462-4afb-a651-9c0410b18b05 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:56.957537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T02:00:00.663355Z digest=sha256:0f23724f02b7b2621f43b1b6c272ac41ffae86880e321d58c96b83b8b975673e

Observation 9645b42f-d3a0-415d-847f-ce765024c3cd · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.392379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T07:17:13.708752Z digest=sha256:bcd5f9ee401a060a1538eb6c72b268923a530660db149f1e529cda728da70043

Observation af7daae2-ef69-4176-a0e9-588baf808caa · inbound

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents cites this paper.

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.650170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T20:16:39.347676Z digest=sha256:ae376da1990a64b2099b4fcc0be9a0f7af6acf9ebfe1d49cf4a15511abf79b5b

Observation 1b4ac855-29fe-4c3f-bd0b-6a6e657fdfe0 · inbound

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training cites this paper.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:e6de513fbb05058624dd44f19f2be5f38d8296cfb49d85eb056a3150b31be046

Observation 42bbac1a-559f-4d9c-b285-f70dfc65cc5b · inbound

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration cites this paper.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.159465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.159465Z digest=sha256:57c5dbed73e6d942d6caa60cdf2e831cd5fbd6062c34e47758c3752468154970

Observation ff2c5ad0-5fd8-464e-8202-baa1d11ad2fd · inbound

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization cites this paper.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.449103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.449103Z digest=sha256:c354cb85d7fa8c7be9b5332beb7227f1b786f8c5bb0be91b75776eeaf154c120