Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:22:19.579220Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2506.05615.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:22:19.579220Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-18T14:02:11.084514Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T14:02:39.803656Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3cfdfe7b-2902-4bd2-861f-3790eca36b4a · outbound
When Maximum Entropy Misleads Policy Optimization write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed8df2e-91d9-4fdd-9e25-d1b2a5a3d09c · outbound
When Maximum Entropy Misleads Policy Optimization Maximum a Posteriori Policy Optimisation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ba359ac-d83a-4881-9cd8-a690560f82f0 · outbound
When Maximum Entropy Misleads Policy Optimization Spinning Up in Deep Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3741c874-6a6b-4ba8-9a4e-6e830276842e · outbound
When Maximum Entropy Misleads Policy Optimization Understanding the impact of entropy on policy optimization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a1d20221-9ce0-42f5-b3f2-fa3e84d00c4b · outbound
When Maximum Entropy Misleads Policy Optimization OpenAI Gym
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb85fd17-36b1-44e7-a447-845bd32a92e0 · outbound
When Maximum Entropy Misleads Policy Optimization Maximum Entropy Reinforcement Learning via Energy-Based Normalizing Flow
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11438e06-5740-4e59-8130-5d739374ddb3 · outbound
When Maximum Entropy Misleads Policy Optimization Maximum Entropy RL (Provably) Solves Some Robust RL Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47be5bd-5272-4dec-9b8b-66325e404ad1 · outbound
When Maximum Entropy Misleads Policy Optimization Taming the Noise in Reinforcement Learning via Soft Updates
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe085a38-ca6f-4a5f-b341-cc86366192f1 · outbound
When Maximum Entropy Misleads Policy Optimization Addressing function approximation error in actor-critic methods
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 322278c8-3405-43bf-90ee-650d5537e22e · outbound
When Maximum Entropy Misleads Policy Optimization Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bb7d50b-58e7-493c-9a8d-38f5ee45f565 · outbound
When Maximum Entropy Misleads Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1a730120-398b-41f2-bbec-73aa4db9edcf · outbound
When Maximum Entropy Misleads Policy Optimization Soft Actor-Critic Algorithms and Applications
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b45be37-f084-425f-91d3-a958e5d58b15 · outbound
When Maximum Entropy Misleads Policy Optimization and Sung, Y
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 175fe72f-766e-4985-857e-00e466ee2af6 · outbound
When Maximum Entropy Misleads Policy Optimization Provably efficient maximum entropy exploration
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8aa36212-437c-4431-b354-5576fe3f3a06 · outbound
When Maximum Entropy Misleads Policy Optimization Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3d1cf22-6f06-4dbe-af98-8544da20c5b7 · outbound
When Maximum Entropy Misleads Policy Optimization Champion-level drone racing using deep reinforcement learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d7dd5da-192f-46b5-b584-f1b184b60e72 · outbound
When Maximum Entropy Misleads Policy Optimization and Sung, Y
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ffd29fb-afab-4ea1-88bb-8879ddb50c78 · outbound
When Maximum Entropy Misleads Policy Optimization Kinematic and dynamic vehicle models for autonomous driving control design
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f5299dd2-2e7d-4417-b210-ba3f39bf8552 · outbound
When Maximum Entropy Misleads Policy Optimization Deep Reinforcement Learning-based UAV Navigation and Control: A Soft Actor-Critic with Hindsight Experience Replay Approach
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ff8de58-e602-4078-a42b-96ae3a2b2693 · outbound
When Maximum Entropy Misleads Policy Optimization Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988004f9-72d7-40bb-a078-d82e59899b90 · outbound
When Maximum Entropy Misleads Policy Optimization Continuous control with deep reinforcement learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59eb3940-8023-4128-b456-7e2ccaae6eb6 · outbound
When Maximum Entropy Misleads Policy Optimization Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 105d893a-77ec-4b90-b5f7-6b126d65be33 · outbound
When Maximum Entropy Misleads Policy Optimization Learning robust perceptive locomotion for quadrupedal robots in the wild
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 080d1d29-62bb-4843-934c-b26275bcb4b1 · outbound
When Maximum Entropy Misleads Policy Optimization Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4daf9d6a-d930-4f42-ae69-d5839c53a0dd · outbound
When Maximum Entropy Misleads Policy Optimization G., D'Souza, J
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7961c1ad-3399-4b34-90e3-82227103d096 · outbound
When Maximum Entropy Misleads Policy Optimization Combining policy gradient and Q-learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e71550c9-1394-4f20-95e2-517adae6e27d · outbound
When Maximum Entropy Misleads Policy Optimization Opencat: Open-source quadruped robot
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17d37bea-62c7-4ccb-8f91-74a0d65efd72 · outbound
When Maximum Entropy Misleads Policy Optimization O., Sedky, A
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5db18f27-50fa-4639-96c0-729c6decf565 · outbound
When Maximum Entropy Misleads Policy Optimization Stable-baselines3: Reliable reinforcement learning implementations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c63c6d64-fe27-4f42-808b-06d2de7a3bb2 · outbound
When Maximum Entropy Misleads Policy Optimization Adversarial Training Can Hurt Generalization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8abdb24a-ff3c-4506-9c82-71e83782a5e8 · outbound
When Maximum Entropy Misleads Policy Optimization Understanding and Mitigating the Tradeoff Between Robustness and Accuracy
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0578b8-ff14-4e66-b2a1-6d8bd692b8c1 · outbound
When Maximum Entropy Misleads Policy Optimization On stochastic optimal control and reinforcement learning by approximate inference
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 91bf7a7a-c3b6-4bf9-86e6-ff230b67743c · outbound
When Maximum Entropy Misleads Policy Optimization A survey of path following control strategies for uavs focused on quadrotors
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a98d043c-5884-4544-9165-a23de4ab58f2 · outbound
When Maximum Entropy Misleads Policy Optimization Trust Region Policy Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ca8698-a3ce-4e2d-a3de-af39c3fbb52f · outbound
When Maximum Entropy Misleads Policy Optimization Proximal Policy Optimization Algorithms
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4759333a-bd67-41eb-9a2e-1ae8992324b4 · outbound
When Maximum Entropy Misleads Policy Optimization M., Vergara, P
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b1ee5d6a-f726-4b09-be7f-bb7e24cc38b9 · outbound
When Maximum Entropy Misleads Policy Optimization Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e5f9703d-9a98-4688-ab6b-7cdd8c8ac3fb · outbound
When Maximum Entropy Misleads Policy Optimization Is robustness the cost of accuracy? – a comprehensive study on the robustness of 18 deep image classification models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9f259164-14cb-49b4-a806-ad92c4130356 · outbound
When Maximum Entropy Misleads Policy Optimization and Karak \"o se, M
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fec21014-f7c5-4cd7-8b51-d30639b59396 · outbound
When Maximum Entropy Misleads Policy Optimization Mujoco: A physics engine for model-based control
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278d8bac-202a-4fff-950e-61e269332467 · outbound
When Maximum Entropy Misleads Policy Optimization Robot trajectory optimization using approximate inference
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fc58c466-bcb2-4fb0-bf7e-757093231711 · outbound
When Maximum Entropy Misleads Policy Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 446f0514-18ab-48ba-84cb-5b9db5ec58ec · outbound
When Maximum Entropy Misleads Policy Optimization Robustness May Be at Odds with Accuracy
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c78c91b-c85e-4c26-90a5-9721dbf43dd2 · outbound
When Maximum Entropy Misleads Policy Optimization Meta-SAC: Auto-tune the Entropy Temperature of Soft Actor-Critic via Metagradient
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f21f248e-f555-4e3d-a345-a0a431a0a37a · outbound
When Maximum Entropy Misleads Policy Optimization Tianshou: a Highly Modularized Deep Reinforcement Learning Library
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a40286e8-707f-43ee-a7bc-6c6d11661dc0 · outbound
When Maximum Entropy Misleads Policy Optimization Karting racing: A revisit to ppo and sac algorithm
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd862ccf-0fa6-4ba0-9a40-62e9f5ef98de · outbound
When Maximum Entropy Misleads Policy Optimization A Closer Look at Accuracy vs. Robustness
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc757881-c934-4bf9-b07c-fa46c66f7b6f · outbound
When Maximum Entropy Misleads Policy Optimization Theoretically principled trade-off between robustness and accuracy
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 376d0516-78de-43d3-a668-855b9b69e303 · outbound
When Maximum Entropy Misleads Policy Optimization Humanoid Parkour Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4fbaa72-699b-40c1-ade6-161f429e7168 · outbound
When Maximum Entropy Misleads Policy Optimization Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 501bdbb6-3a21-4c66-8c99-9df64c8a0f33 · outbound
When Maximum Entropy Misleads Policy Optimization D., Maas, A
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 56f51a5b-be0c-4dca-992f-1a07088813b8 · outbound
When Maximum Entropy Misleads Policy Optimization D., Bagnell, D., and Dey, A
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e5ebc49-9234-42de-9a5e-aea75042f0fd · inbound
Failure Modes of Maximum Entropy RLHF When Maximum Entropy Misleads Policy Optimization
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f247a322-b1c6-4be2-9a99-b97a6a08b6c7 · inbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization When Maximum Entropy Misleads Policy Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.