Pith. sign in

Paper Citation Record · LEDGER

When Maximum Entropy Misleads Policy Optimization

As of 10 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2506.05615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05615 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:22:19.579220Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T14:02:11.084514Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T14:02:39.803656Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cfdfe7b-2902-4bd2-861f-3790eca36b4a · outbound

This paper cites write newline.

When Maximum Entropy Misleads Policy Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:16.719555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:16.719555Z digest=sha256:b8dd60d83a830e8c19f33ace98295f2cf7b4fcc4593cf2efe5a9fef7f0333181

Observation 7ed8df2e-91d9-4fdd-9e25-d1b2a5a3d09c · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

When Maximum Entropy Misleads Policy Optimization Maximum a Posteriori Policy Optimisation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:16.780892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:16.780892Z digest=sha256:bce1164e52b8da70fc17b5fa9a4fee998376717fa3d7deda5a09cd63d74c2bb1

Observation 3ba359ac-d83a-4881-9cd8-a690560f82f0 · outbound

This paper cites Spinning Up in Deep Reinforcement Learning.

When Maximum Entropy Misleads Policy Optimization Spinning Up in Deep Reinforcement Learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.957240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:16.828271Z digest=sha256:db23d8eb28373b3f29f154530883073352aabbb6c069de55260c23fd487d27b5

Observation 3741c874-6a6b-4ba8-9a4e-6e830276842e · outbound

This paper cites Understanding the impact of entropy on policy optimization.

When Maximum Entropy Misleads Policy Optimization Understanding the impact of entropy on policy optimization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.949078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:16.877406Z digest=sha256:488d3f28de60c33e37315e263502c8e0bf686365d9ace8fddb591bb06eed2631

Observation a1d20221-9ce0-42f5-b3f2-fa3e84d00c4b · outbound

This paper cites OpenAI Gym.

When Maximum Entropy Misleads Policy Optimization OpenAI Gym

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:16.921856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:16.921856Z digest=sha256:305dee57e531348d923899d9de67fdb14838fd8cb468feb5d443c0bccd5c8888

Observation fb85fd17-36b1-44e7-a447-845bd32a92e0 · outbound

This paper cites Maximum Entropy Reinforcement Learning via Energy-Based Normalizing Flow.

When Maximum Entropy Misleads Policy Optimization Maximum Entropy Reinforcement Learning via Energy-Based Normalizing Flow

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:16.992095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:16.992095Z digest=sha256:96071d96b9aa1e57f31aaff98ca66e58e97fb3f2551c8a2e5d7811f932f4874b

Observation 11438e06-5740-4e59-8130-5d739374ddb3 · outbound

This paper cites Maximum Entropy RL (Provably) Solves Some Robust RL Problems.

When Maximum Entropy Misleads Policy Optimization Maximum Entropy RL (Provably) Solves Some Robust RL Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.058715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.058715Z digest=sha256:22a7ac6bd4f67bc4c3c1d71bfff7e7ebf08f5c847bbf6f69f25d884686bf6464

Observation a47be5bd-5272-4dec-9b8b-66325e404ad1 · outbound

This paper cites Taming the Noise in Reinforcement Learning via Soft Updates.

When Maximum Entropy Misleads Policy Optimization Taming the Noise in Reinforcement Learning via Soft Updates

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.125983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.125983Z digest=sha256:f5558a819700e2af7e52768ba50c9b259553619e65ca3fb31a13634fed1332fb

Observation fe085a38-ca6f-4a5f-b341-cc86366192f1 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

When Maximum Entropy Misleads Policy Optimization Addressing function approximation error in actor-critic methods

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.176212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.176212Z digest=sha256:f69ed0a2906f61744f5ee13fe39e050bf273847e22b782ac3f5f37467243a7dc

Observation 322278c8-3405-43bf-90ee-650d5537e22e · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:19.937401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.221044Z digest=sha256:3e732eff9253a575fc992b23e6ca063aa5674666732ae0557ee3b96e420ee7fc

Observation 3bb7d50b-58e7-493c-9a8d-38f5ee45f565 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

When Maximum Entropy Misleads Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.930355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.270932Z digest=sha256:0588f0ded03989cbb500d891221910388e99eddc4d2310fe4128290dd54a50b6

Observation 1a730120-398b-41f2-bbec-73aa4db9edcf · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

When Maximum Entropy Misleads Policy Optimization Soft Actor-Critic Algorithms and Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.317512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.317512Z digest=sha256:262d98936cee4b728cd4e82582898e6aa6e5e0ef766afb9d11108eea1543c6e5

Observation 5b45be37-f084-425f-91d3-a958e5d58b15 · outbound

This paper cites and Sung, Y.

When Maximum Entropy Misleads Policy Optimization and Sung, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.922589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.388420Z digest=sha256:dc131aa43d635959bb43f4418fc5c52db4f7aa2f3a7f374cd583438c6b6e9581

Observation 175fe72f-766e-4985-857e-00e466ee2af6 · outbound

This paper cites Provably efficient maximum entropy exploration.

When Maximum Entropy Misleads Policy Optimization Provably efficient maximum entropy exploration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.914968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.464584Z digest=sha256:a116d0b931453a9df3b84e86b3898c3516a470acd4be8049ba36169df344f2b4

Observation 8aa36212-437c-4431-b354-5576fe3f3a06 · outbound

This paper cites Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning.

When Maximum Entropy Misleads Policy Optimization Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.490373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.490373Z digest=sha256:b1147314766a1afa7bf2b866569dac077d989f4cb295eee55f3657a656219db8

Observation a3d1cf22-6f06-4dbe-af98-8544da20c5b7 · outbound

This paper cites Champion-level drone racing using deep reinforcement learning.

When Maximum Entropy Misleads Policy Optimization Champion-level drone racing using deep reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.907705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.593655Z digest=sha256:7802603bf41b463943f545eb3b34d749f8656334b1b29cb6c74d6b1dbd85066a

Observation 9d7dd5da-192f-46b5-b584-f1b184b60e72 · outbound

This paper cites and Sung, Y.

When Maximum Entropy Misleads Policy Optimization and Sung, Y

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.900176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.744302Z digest=sha256:0ff2232fefd7f433bb6e853e8dc5cbc229647a0d7740f01dd2dd944e53a7c7bb

Observation 1ffd29fb-afab-4ea1-88bb-8879ddb50c78 · outbound

This paper cites Kinematic and dynamic vehicle models for autonomous driving control design.

When Maximum Entropy Misleads Policy Optimization Kinematic and dynamic vehicle models for autonomous driving control design

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.892111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.831416Z digest=sha256:2e79f42cbd42b2a6538b66e1e48a917bb373c326c68858b1a2d6bdfaf1f17f3b

Observation f5299dd2-2e7d-4417-b210-ba3f39bf8552 · outbound

This paper cites Deep Reinforcement Learning-based UAV Navigation and Control: A Soft Actor-Critic with Hindsight Experience Replay Approach.

When Maximum Entropy Misleads Policy Optimization Deep Reinforcement Learning-based UAV Navigation and Control: A Soft Actor-Critic with Hindsight Experience Replay Approach

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:22:19.684788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.963354Z digest=sha256:e9e7948c2aebae01e26e5717b917ee148efe26b618fb1b92908668b604360250

Observation 7ff8de58-e602-4078-a42b-96ae3a2b2693 · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

When Maximum Entropy Misleads Policy Optimization Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:18.076144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:18.076144Z digest=sha256:c7b0c25acfdbea15da0b0c769f36da313aceded7c61e020bb75030c03d3de55e

Observation 988004f9-72d7-40bb-a078-d82e59899b90 · outbound

This paper cites Continuous control with deep reinforcement learning.

When Maximum Entropy Misleads Policy Optimization Continuous control with deep reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:18.196750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:18.196750Z digest=sha256:04c524ab03b8b7563b201143aaa40ce0c9de0f4ff4462139b94f52976a242132

Observation 59eb3940-8023-4128-b456-7e2ccaae6eb6 · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:19.883223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.282718Z digest=sha256:cfc0e39190afc482a6a95563c9e2c228da2f23c9481d5eac4535122fb691f547

Observation 105d893a-77ec-4b90-b5f7-6b126d65be33 · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

When Maximum Entropy Misleads Policy Optimization Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.875677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.396483Z digest=sha256:7bcb9f67a35917a46085a823931c72c38dab047565f870d2b0f9dc638531f66a

Observation 080d1d29-62bb-4843-934c-b26275bcb4b1 · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:19.867591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.554171Z digest=sha256:6c16799a44dac6c1be46203c041d6a03403eafe1d878c0912e114d11c7c33e93

Observation 4daf9d6a-d930-4f42-ae69-d5839c53a0dd · outbound

This paper cites G., D'Souza, J.

When Maximum Entropy Misleads Policy Optimization G., D'Souza, J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.859215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.668358Z digest=sha256:20e592e1776d146d9a3a53599a1ac5f6e07bb0de9005526b47bb3efa7c3e7c56

Observation 7961c1ad-3399-4b34-90e3-82227103d096 · outbound

This paper cites Combining policy gradient and Q-learning.

When Maximum Entropy Misleads Policy Optimization Combining policy gradient and Q-learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:18.826187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:18.826187Z digest=sha256:3a8f0e51ae704855facaedad9d3228481225887d76844d7d70fc41a5035fd4c9

Observation e71550c9-1394-4f20-95e2-517adae6e27d · outbound

This paper cites Opencat: Open-source quadruped robot.

When Maximum Entropy Misleads Policy Optimization Opencat: Open-source quadruped robot

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.851332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.944410Z digest=sha256:e9bc009edc555fb5caf8198c99704d53649ae796b6585c9130ebef4903e8d21b

Observation 17d37bea-62c7-4ccb-8f91-74a0d65efd72 · outbound

This paper cites O., Sedky, A.

When Maximum Entropy Misleads Policy Optimization O., Sedky, A

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.843407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.048347Z digest=sha256:d77f9b4f0945a7292faf229de68d99bedf5cbe47e1e02919def01dadd62d2a5a

Observation 5db18f27-50fa-4639-96c0-729c6decf565 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

When Maximum Entropy Misleads Policy Optimization Stable-baselines3: Reliable reinforcement learning implementations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.834678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.117262Z digest=sha256:9546c2745e60b2aacd6f7ccdbbc85154cd4d40ff1f6038775337c880f11f70c6

Observation c63c6d64-fe27-4f42-808b-06d2de7a3bb2 · outbound

This paper cites Adversarial Training Can Hurt Generalization.

When Maximum Entropy Misleads Policy Optimization Adversarial Training Can Hurt Generalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.229507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.229507Z digest=sha256:ec0c3f51b27dfe3abf8d78f68786f6f1fec4c860c3aab09399ce5c3d951502d0

Observation 8abdb24a-ff3c-4506-9c82-71e83782a5e8 · outbound

This paper cites Understanding and Mitigating the Tradeoff Between Robustness and Accuracy.

When Maximum Entropy Misleads Policy Optimization Understanding and Mitigating the Tradeoff Between Robustness and Accuracy

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.326777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.326777Z digest=sha256:672b0b330cb3548935087d0e025d906f96ffa30072c70aa19bacfe6e088d02b2

Observation ba0578b8-ff14-4e66-b2a1-6d8bd692b8c1 · outbound

This paper cites On stochastic optimal control and reinforcement learning by approximate inference.

When Maximum Entropy Misleads Policy Optimization On stochastic optimal control and reinforcement learning by approximate inference

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.826928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.443620Z digest=sha256:cb043a6910c7eb04684113423ca7fc03ada09e2d10f03b6ee7191a2ce3bee60d

Observation 91bf7a7a-c3b6-4bf9-86e6-ff230b67743c · outbound

This paper cites A survey of path following control strategies for uavs focused on quadrotors.

When Maximum Entropy Misleads Policy Optimization A survey of path following control strategies for uavs focused on quadrotors

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.818218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.533147Z digest=sha256:99740e9ec4e3f17712e14670669978ab2a3a7fa8ead8f92f1dcc0d91d9c80ce2

Observation a98d043c-5884-4544-9165-a23de4ab58f2 · outbound

This paper cites Trust Region Policy Optimization.

When Maximum Entropy Misleads Policy Optimization Trust Region Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.535372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.535372Z digest=sha256:9e2236f689dabc05ca38499779949b925d761002dc1488399fafaf8179905b4a

Observation 68ca8698-a3ce-4e2d-a3de-af39c3fbb52f · outbound

This paper cites Proximal Policy Optimization Algorithms.

When Maximum Entropy Misleads Policy Optimization Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.537569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.537569Z digest=sha256:1409d40a0b9c85ca23226c7dcf93f615eb6c079072dc2627bee82c0e7a504db0

Observation 4759333a-bd67-41eb-9a2e-1ae8992324b4 · outbound

This paper cites M., Vergara, P.

When Maximum Entropy Misleads Policy Optimization M., Vergara, P

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.810564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.541152Z digest=sha256:4297e13cb31b98020dca9f61496bae05ddea6aa45c812965997edb9dbe31a323

Observation b1ee5d6a-f726-4b09-be7f-bb7e24cc38b9 · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:19.802698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.543450Z digest=sha256:34478b35e6950eb050e9ceaa55d776265b08beba490b4d490e720ac0aa09dbca

Observation e5f9703d-9a98-4688-ab6b-7cdd8c8ac3fb · outbound

This paper cites Is robustness the cost of accuracy? – a comprehensive study on the robustness of 18 deep image classification models.

When Maximum Entropy Misleads Policy Optimization Is robustness the cost of accuracy? – a comprehensive study on the robustness of 18 deep image classification models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.794990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.545689Z digest=sha256:33aa42e0b1673da68ad3625cd79721d8138b6b007c0ca28a4c5ab879756fceb2

Observation 9f259164-14cb-49b4-a806-ad92c4130356 · outbound

This paper cites and Karak \"o se, M.

When Maximum Entropy Misleads Policy Optimization and Karak \"o se, M

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.785849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.547605Z digest=sha256:140f49ec5c6b56d2b70981e99ab5c3d4d374452155085691fa9fb48fcaa0ea7d

Observation fec21014-f7c5-4cd7-8b51-d30639b59396 · outbound

This paper cites Mujoco: A physics engine for model-based control.

When Maximum Entropy Misleads Policy Optimization Mujoco: A physics engine for model-based control

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.549765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.549765Z digest=sha256:6b051e572dab3a5af218bc2458a7a5f23b7c0d99e88f89388667eb3f3869c249

Observation 278d8bac-202a-4fff-950e-61e269332467 · outbound

This paper cites Robot trajectory optimization using approximate inference.

When Maximum Entropy Misleads Policy Optimization Robot trajectory optimization using approximate inference

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.773741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.552457Z digest=sha256:c88e634663259f1164db23ef957921ece61173a6cb7fa66900decfd5d7ac9e98

Observation fc58c466-bcb2-4fb0-bf7e-757093231711 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

When Maximum Entropy Misleads Policy Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.555066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.555066Z digest=sha256:6d6cebb5c8cbe10ca127222a4bdb6bf3c3aa632c962e674f77f223e97b227fae

Observation 446f0514-18ab-48ba-84cb-5b9db5ec58ec · outbound

This paper cites Robustness May Be at Odds with Accuracy.

When Maximum Entropy Misleads Policy Optimization Robustness May Be at Odds with Accuracy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.557477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.557477Z digest=sha256:289415e1073adfd2e8c607e34ac3f6ba97cae99fc6d0c4846aa812db01b3ac1c

Observation 2c78c91b-c85e-4c26-90a5-9721dbf43dd2 · outbound

This paper cites Meta-SAC: Auto-tune the Entropy Temperature of Soft Actor-Critic via Metagradient.

When Maximum Entropy Misleads Policy Optimization Meta-SAC: Auto-tune the Entropy Temperature of Soft Actor-Critic via Metagradient

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.560314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.560314Z digest=sha256:b407ffabf2ef60895958e805e34c2d1206c0942d0ef9fd3abba24d61ef6398e0

Observation f21f248e-f555-4e3d-a345-a0a431a0a37a · outbound

This paper cites Tianshou: a Highly Modularized Deep Reinforcement Learning Library.

When Maximum Entropy Misleads Policy Optimization Tianshou: a Highly Modularized Deep Reinforcement Learning Library

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.562842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.562842Z digest=sha256:14a2c587b0559c27e9a8fd2725f21532eacd49e17c56be3978a3a10efeacb73d

Observation a40286e8-707f-43ee-a7bc-6c6d11661dc0 · outbound

This paper cites Karting racing: A revisit to ppo and sac algorithm.

When Maximum Entropy Misleads Policy Optimization Karting racing: A revisit to ppo and sac algorithm

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.765245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.565395Z digest=sha256:6a616b3fd99b98c2a63edd82962380059c4a8f398f5972706ad234bc4dc26861

Observation bd862ccf-0fa6-4ba0-9a40-62e9f5ef98de · outbound

This paper cites A Closer Look at Accuracy vs. Robustness.

When Maximum Entropy Misleads Policy Optimization A Closer Look at Accuracy vs. Robustness

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.567927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.567927Z digest=sha256:03e623172a6dd1020035a932a9901473c90f5ebc333efe325924de38fc3f2fe7

Observation dc757881-c934-4bf9-b07c-fa46c66f7b6f · outbound

This paper cites Theoretically principled trade-off between robustness and accuracy.

When Maximum Entropy Misleads Policy Optimization Theoretically principled trade-off between robustness and accuracy

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.757149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.570121Z digest=sha256:ef019f1fc813e6a627b9963dd024961b2518a4920353db803155ceb1e656f7f7

Observation 376d0516-78de-43d3-a668-855b9b69e303 · outbound

This paper cites Humanoid Parkour Learning.

When Maximum Entropy Misleads Policy Optimization Humanoid Parkour Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.572182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.572182Z digest=sha256:174fd5fae33f8605ba8d729aec3174d393fa946fa532c31f0f59950633c6b313

Observation d4fbaa72-699b-40c1-ade6-161f429e7168 · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.574355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.574355Z digest=sha256:eeef107c53283ec7a40ea03a3c1e57ccd588a4c16458deaa2dbc9b4d005c7225

Observation 501bdbb6-3a21-4c66-8c99-9df64c8a0f33 · outbound

This paper cites D., Maas, A.

When Maximum Entropy Misleads Policy Optimization D., Maas, A

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.745549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.576834Z digest=sha256:0d0b299dd4cee2d12676b99ed46c0732155674ce160b0093f0701f65a46bb200

Observation 56f51a5b-be0c-4dca-992f-1a07088813b8 · outbound

This paper cites D., Bagnell, D., and Dey, A.

When Maximum Entropy Misleads Policy Optimization D., Bagnell, D., and Dey, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.736588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.579220Z digest=sha256:4eeadcca331044d51ad92ff3b1de060b9b7c4fc0e0f07f32a5926034761c8df9

Pith citing papers

Observation 0e5ebc49-9234-42de-9a5e-aea75042f0fd · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF When Maximum Entropy Misleads Policy Optimization

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.807781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:ca920121650322d52368f27e797fb11d04e67e5725ce280d7ae8cfe804e3e9fa

Observation f247a322-b1c6-4be2-9a99-b97a6a08b6c7 · inbound

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization cites this paper.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization When Maximum Entropy Misleads Policy Optimization

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.858360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:9910dd4314c6837b8f984e9e5d79f52e7166d20172ba1e23e9837a15c0ceeea1