Pith. sign in

Paper Citation Record · LEDGER

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization

As of 21 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2505.17714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17714 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:45:01.725671Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed48a1ba-0985-4002-8884-99638701bedf · outbound

This paper cites Human-level control through deep reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Human-level control through deep reinforcement learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:59.795938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:59.795938Z digest=sha256:d14eb710023801c3e504fde5ce476bbf7a01cbecef159f8673925fad37a8bd12

Observation ba071adb-7693-41f6-b689-43307b421dc9 · outbound

This paper cites Benchmarking deep reinforcement learning for continuous control,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Benchmarking deep reinforcement learning for continuous control,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:06.440862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:44:59.896292Z digest=sha256:79f6e8b3552c47d68a4572519ee113397d892ac3339ebf7de697f5548db00553

Observation 9fa4db05-1286-48d5-982b-b0542b980360 · outbound

This paper cites Proximal Policy Optimization Algorithms.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Proximal Policy Optimization Algorithms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:00.026191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:00.026191Z digest=sha256:a27b5860820272dd85f8123b97ef37bbb78d8921484d499afa5de3fe02ada101

Observation aaea74fd-0dfa-4cf0-80f6-8ce220d945b5 · outbound

This paper cites Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:06.094173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.118894Z digest=sha256:b59634b9cbac895d88070c0686c315174e9f4ee58630101dec253a9f8e3bc8dd

Observation af84fa71-8172-418a-a21f-51ccf5973597 · outbound

This paper cites Trust region policy optimization,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Trust region policy optimization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:05.758067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.173993Z digest=sha256:4560c1eaac8e0146aa59ff3bee07218b18ba859342fe956df3bb85bd7dd688e3

Observation 1ea400c7-b34f-496a-8f83-5b92ddac29bc · outbound

This paper cites Annealed policy optimization for deep reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Annealed policy optimization for deep reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:05.421447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.226610Z digest=sha256:0fc10de055401ebdc89c1a11483ca6e6cb203dc4511af2a6af67c39a6e58f612

Observation e826bed6-4128-40b3-8976-794773bad0f1 · outbound

This paper cites Understanding the impact of entropy on policy optimization,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Understanding the impact of entropy on policy optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:05.066997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.328576Z digest=sha256:ae0c2a82d92f652ad76b0a74d3397b94b62582da2f702d0e495bf9bec5ce6dc8

Observation 9f4fba45-8f02-4dd0-a811-d838d83d1529 · outbound

This paper cites Normalized policy gradients for reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Normalized policy gradients for reinforcement learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:04.720748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.425882Z digest=sha256:1947b874fa4b4ad8c6bc7db1951ccffcdb2180e8324a7be7553812cc8235f822

Observation d435a8a6-a97a-4815-8819-1cf31dee32e2 · outbound

This paper cites Sutton and A.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Sutton and A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:04.445147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.489904Z digest=sha256:fde0ab1de56c0f9d36de0a4785b5e7d1603d4e2d7a1d2c662a6c85cfc1ee9ca0

Observation db774f95-8495-4d18-87a8-156d86dc4837 · outbound

This paper cites Reinforcement learning: A survey,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Reinforcement learning: A survey,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:04.108939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.593465Z digest=sha256:cbf26dd7ea0ab7ecad657afadcb15218da48ed075b33a06854e19a0d4a2fe96c

Observation 49f59f67-a684-43b3-9e07-d88295652ad2 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Simple statistical gradient-following algorithms for connectionist reinforcement learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:03.861939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.667690Z digest=sha256:00795fb872f4fa24ef99f4a51db5fbf2c1e14862f9d5c821a54e550a80968fe1

Observation ce536db6-f0aa-47ec-bbae-acfece1719f7 · outbound

This paper cites High-dimensional continuous control using generalized advantage estimation,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization High-dimensional continuous control using generalized advantage estimation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:03.634034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.768484Z digest=sha256:89a334ae2eb9ca8f71a3812dcfbed36835c1ee08980eee7e885eb5feb654bb84

Observation 350014f5-19c8-4dc2-b356-18a77b7a37c4 · outbound

This paper cites Grandmaster level in StarCraft II using multi-agent reinforcement learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Grandmaster level in StarCraft II using multi-agent reinforcement learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:03.466030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:00.891898Z digest=sha256:268c0ca878c7510ed565dbcf55896d36a63eb20fb8e02332b33efabcd754b400

Observation 7dc67ac9-c92c-44ff-88af-46b978cdf3e7 · outbound

This paper cites Annealed policy optimization,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Annealed policy optimization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:03.117391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:01.006215Z digest=sha256:d436f0f35eae09675358e432f69adb0098284be998c911b65c6192c96a328561

Observation f77c92cb-f40a-42c5-b9bd-a3ef5461038d · outbound

This paper cites Noisy networks for exploration,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Noisy networks for exploration,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:02.878463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:01.112097Z digest=sha256:e56349e214e072bf998cb1f31ead7350d1157da35a36c3e6afc5a09241126f75

Observation 725dfd50-53b0-4a33-abe4-67a6f5dbd49b · outbound

This paper cites Deep Reinforcement Learning: An Overview,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Deep Reinforcement Learning: An Overview,

Reference 16

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:45:02.226161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:01.223279Z digest=sha256:7f38ee57d450fcfb71f3467c3de929a3a2def4e9b21106288f103d39eefe9b31

Observation 5a5b95cf-715c-4aaf-8a2d-a4aa84a5f2d4 · outbound

This paper cites Constrained Policy Optimization,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Constrained Policy Optimization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:02.628187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:01.373170Z digest=sha256:72341334bbdbc7366b6411ffe3e31e907873b27327cab51f95002ad100987e49

Observation 6e852697-3be3-4650-a1f9-4a74d4c7c120 · outbound

This paper cites A Study on Overfitting in Deep Reinforcement Learning,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization A Study on Overfitting in Deep Reinforcement Learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:02.404638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:45:01.506161Z digest=sha256:c329465180931932bab2d6874175395b6935662de6fc5eb53e042e2be3227a47

Observation b155d807-98c2-4b2d-967a-679e0e6b73db · outbound

This paper cites Optimizing Customer Satisfaction Through Sentiment Analysis: A BERT-Based Machine Learning Approach to Extract Insights,.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Optimizing Customer Satisfaction Through Sentiment Analysis: A BERT-Based Machine Learning Approach to Extract Insights,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:01.598606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:01.598606Z digest=sha256:d941d911532cf33698a1b3395845e4b142ce4653139698ec64fba62c64320c46

Observation ec1f8a1a-a535-4f03-9ffd-7d932f0ab305 · outbound

This paper cites Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications.

PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:01.725671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:01.725671Z digest=sha256:d12c74d8966c7970c80bef688856102896a3b165d0df4f2c75ab17fd886bbcac

Pith citing papers

No inbound Pith citation observations are available.