Pith. sign in

Paper Citation Record · LEDGER

When Maximum Entropy Misleads Policy Optimization

As of 10 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2506.05615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05615 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:22:19.579220Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T14:02:11.084514Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T14:02:39.803656Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cfdfe7b-2902-4bd2-861f-3790eca36b4a · outbound

This paper cites write newline.

When Maximum Entropy Misleads Policy Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:16.719555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:16.719555Z digest=sha256:b8dd60d83a830e8c19f33ace98295f2cf7b4fcc4593cf2efe5a9fef7f0333181

Observation 7ed8df2e-91d9-4fdd-9e25-d1b2a5a3d09c · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

When Maximum Entropy Misleads Policy Optimization Maximum a Posteriori Policy Optimisation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:16.780892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:16.780892Z digest=sha256:bce1164e52b8da70fc17b5fa9a4fee998376717fa3d7deda5a09cd63d74c2bb1

Observation 3ba359ac-d83a-4881-9cd8-a690560f82f0 · outbound

This paper cites Spinning Up in Deep Reinforcement Learning.

When Maximum Entropy Misleads Policy Optimization Spinning Up in Deep Reinforcement Learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.957240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:16.828271Z digest=sha256:2510882604c629409207da141cbb0b993a01a371d1c54eed7d6cba080cb87284

Observation 3741c874-6a6b-4ba8-9a4e-6e830276842e · outbound

This paper cites Understanding the impact of entropy on policy optimization.

When Maximum Entropy Misleads Policy Optimization Understanding the impact of entropy on policy optimization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.949078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:16.877406Z digest=sha256:e665546caa217c4bad75cf90026bad9c58f049e4f74f46f3b78fbf9593534924

Observation a1d20221-9ce0-42f5-b3f2-fa3e84d00c4b · outbound

This paper cites OpenAI Gym.

When Maximum Entropy Misleads Policy Optimization OpenAI Gym

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:16.921856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:16.921856Z digest=sha256:305dee57e531348d923899d9de67fdb14838fd8cb468feb5d443c0bccd5c8888

Observation fb85fd17-36b1-44e7-a447-845bd32a92e0 · outbound

This paper cites Maximum Entropy Reinforcement Learning via Energy-Based Normalizing Flow.

When Maximum Entropy Misleads Policy Optimization Maximum Entropy Reinforcement Learning via Energy-Based Normalizing Flow

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:16.992095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:16.992095Z digest=sha256:96071d96b9aa1e57f31aaff98ca66e58e97fb3f2551c8a2e5d7811f932f4874b

Observation 11438e06-5740-4e59-8130-5d739374ddb3 · outbound

This paper cites Maximum Entropy RL (Provably) Solves Some Robust RL Problems.

When Maximum Entropy Misleads Policy Optimization Maximum Entropy RL (Provably) Solves Some Robust RL Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.058715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.058715Z digest=sha256:22a7ac6bd4f67bc4c3c1d71bfff7e7ebf08f5c847bbf6f69f25d884686bf6464

Observation a47be5bd-5272-4dec-9b8b-66325e404ad1 · outbound

This paper cites Taming the Noise in Reinforcement Learning via Soft Updates.

When Maximum Entropy Misleads Policy Optimization Taming the Noise in Reinforcement Learning via Soft Updates

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.125983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.125983Z digest=sha256:f5558a819700e2af7e52768ba50c9b259553619e65ca3fb31a13634fed1332fb

Observation fe085a38-ca6f-4a5f-b341-cc86366192f1 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

When Maximum Entropy Misleads Policy Optimization Addressing function approximation error in actor-critic methods

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.176212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.176212Z digest=sha256:f69ed0a2906f61744f5ee13fe39e050bf273847e22b782ac3f5f37467243a7dc

Observation 322278c8-3405-43bf-90ee-650d5537e22e · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:19.937401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.221044Z digest=sha256:307ff6105180dfb411bc8873525546275ce6501c4f4904da05ad477ee0f38505

Observation 3bb7d50b-58e7-493c-9a8d-38f5ee45f565 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

When Maximum Entropy Misleads Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.930355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.270932Z digest=sha256:12e4d56da962b3e148d26b6c5a88a341d26309cafc1fe2f56cde277c414c7d42

Observation 1a730120-398b-41f2-bbec-73aa4db9edcf · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

When Maximum Entropy Misleads Policy Optimization Soft Actor-Critic Algorithms and Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.317512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.317512Z digest=sha256:262d98936cee4b728cd4e82582898e6aa6e5e0ef766afb9d11108eea1543c6e5

Observation 5b45be37-f084-425f-91d3-a958e5d58b15 · outbound

This paper cites and Sung, Y.

When Maximum Entropy Misleads Policy Optimization and Sung, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.922589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.388420Z digest=sha256:f29cc816912fcf00d6fd199139f57171cffbd66cd04d01ac514c3315c6b32d09

Observation 175fe72f-766e-4985-857e-00e466ee2af6 · outbound

This paper cites Provably efficient maximum entropy exploration.

When Maximum Entropy Misleads Policy Optimization Provably efficient maximum entropy exploration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.914968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.464584Z digest=sha256:882185a2316760f7f223ba90914d8b522e93c1f1ebbb91f7f745bab69fa3d065

Observation 8aa36212-437c-4431-b354-5576fe3f3a06 · outbound

This paper cites Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning.

When Maximum Entropy Misleads Policy Optimization Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:17.490373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:17.490373Z digest=sha256:b1147314766a1afa7bf2b866569dac077d989f4cb295eee55f3657a656219db8

Observation a3d1cf22-6f06-4dbe-af98-8544da20c5b7 · outbound

This paper cites Champion-level drone racing using deep reinforcement learning.

When Maximum Entropy Misleads Policy Optimization Champion-level drone racing using deep reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.907705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.593655Z digest=sha256:3f77cc95ffae6f1c9f81221ec1be01cbdb1dc01d5ac6411cb677982674bee9f3

Observation 9d7dd5da-192f-46b5-b584-f1b184b60e72 · outbound

This paper cites and Sung, Y.

When Maximum Entropy Misleads Policy Optimization and Sung, Y

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.900176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.744302Z digest=sha256:8ad0d8f7d453762f8c885dccadfb8a13e47d01c2cfe320cd588ef1c33633b66b

Observation 1ffd29fb-afab-4ea1-88bb-8879ddb50c78 · outbound

This paper cites Kinematic and dynamic vehicle models for autonomous driving control design.

When Maximum Entropy Misleads Policy Optimization Kinematic and dynamic vehicle models for autonomous driving control design

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.892111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.831416Z digest=sha256:14c374fe03dc72e5c8ec4aad7c0a95ba96cd257028af7e967366052c8d54140b

Observation f5299dd2-2e7d-4417-b210-ba3f39bf8552 · outbound

This paper cites Deep Reinforcement Learning-based UAV Navigation and Control: A Soft Actor-Critic with Hindsight Experience Replay Approach.

When Maximum Entropy Misleads Policy Optimization Deep Reinforcement Learning-based UAV Navigation and Control: A Soft Actor-Critic with Hindsight Experience Replay Approach

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:22:19.684788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:17.963354Z digest=sha256:b1541f539dcd05c6063d368f3230e522bd4471e942d70e17571dd30703175259

Observation 7ff8de58-e602-4078-a42b-96ae3a2b2693 · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

When Maximum Entropy Misleads Policy Optimization Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:18.076144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:18.076144Z digest=sha256:c7b0c25acfdbea15da0b0c769f36da313aceded7c61e020bb75030c03d3de55e

Observation 988004f9-72d7-40bb-a078-d82e59899b90 · outbound

This paper cites Continuous control with deep reinforcement learning.

When Maximum Entropy Misleads Policy Optimization Continuous control with deep reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:18.196750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:18.196750Z digest=sha256:04c524ab03b8b7563b201143aaa40ce0c9de0f4ff4462139b94f52976a242132

Observation 59eb3940-8023-4128-b456-7e2ccaae6eb6 · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:19.883223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.282718Z digest=sha256:64cc07fc61cf8d57ebad50b387dc1e5f5c804be9bb40bc9f3d125d6a2a89dbb2

Observation 105d893a-77ec-4b90-b5f7-6b126d65be33 · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

When Maximum Entropy Misleads Policy Optimization Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.875677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.396483Z digest=sha256:2710c80dfded8ef6d45dde2f9cbfbdb576376892f04bc9b592a79aeb768e9e05

Observation 080d1d29-62bb-4843-934c-b26275bcb4b1 · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:19.867591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.554171Z digest=sha256:d213d5618b23a15a15e391e855a24decd48d4eb0087a5169194969039ced5280

Observation 4daf9d6a-d930-4f42-ae69-d5839c53a0dd · outbound

This paper cites G., D'Souza, J.

When Maximum Entropy Misleads Policy Optimization G., D'Souza, J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.859215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.668358Z digest=sha256:a18daa870c39491a96d96c5b570819b2273fb971e832a4199b3b96afcd8719d7

Observation 7961c1ad-3399-4b34-90e3-82227103d096 · outbound

This paper cites Combining policy gradient and Q-learning.

When Maximum Entropy Misleads Policy Optimization Combining policy gradient and Q-learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:18.826187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:18.826187Z digest=sha256:0f4780616f3afe323d21f5e69f68da2cb81b0ca11be322f1abf1fb0d368a3346

Observation e71550c9-1394-4f20-95e2-517adae6e27d · outbound

This paper cites Opencat: Open-source quadruped robot.

When Maximum Entropy Misleads Policy Optimization Opencat: Open-source quadruped robot

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.851332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:18.944410Z digest=sha256:518266a09d927385cb27b3560a3c64dff758f1d77cc51287126c05840ca0645a

Observation 17d37bea-62c7-4ccb-8f91-74a0d65efd72 · outbound

This paper cites O., Sedky, A.

When Maximum Entropy Misleads Policy Optimization O., Sedky, A

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.843407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.048347Z digest=sha256:6a5ea1d88ac50cfea889b421d271296af1f91369d607c590cce75cec3d58c545

Observation 5db18f27-50fa-4639-96c0-729c6decf565 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

When Maximum Entropy Misleads Policy Optimization Stable-baselines3: Reliable reinforcement learning implementations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.834678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.117262Z digest=sha256:da8062a829a6bcfe0bbbb05547d87a743e422b8c4fc1209f7fefc239dc190a64

Observation c63c6d64-fe27-4f42-808b-06d2de7a3bb2 · outbound

This paper cites Adversarial Training Can Hurt Generalization.

When Maximum Entropy Misleads Policy Optimization Adversarial Training Can Hurt Generalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.229507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.229507Z digest=sha256:ec0c3f51b27dfe3abf8d78f68786f6f1fec4c860c3aab09399ce5c3d951502d0

Observation 8abdb24a-ff3c-4506-9c82-71e83782a5e8 · outbound

This paper cites Understanding and Mitigating the Tradeoff Between Robustness and Accuracy.

When Maximum Entropy Misleads Policy Optimization Understanding and Mitigating the Tradeoff Between Robustness and Accuracy

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.326777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.326777Z digest=sha256:672b0b330cb3548935087d0e025d906f96ffa30072c70aa19bacfe6e088d02b2

Observation ba0578b8-ff14-4e66-b2a1-6d8bd692b8c1 · outbound

This paper cites On stochastic optimal control and reinforcement learning by approximate inference.

When Maximum Entropy Misleads Policy Optimization On stochastic optimal control and reinforcement learning by approximate inference

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.826928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.443620Z digest=sha256:debcc12771a5ad3f46e9dcffd7f4bbceea5110a6691000d3fca7e90ea6d337ac

Observation 91bf7a7a-c3b6-4bf9-86e6-ff230b67743c · outbound

This paper cites A survey of path following control strategies for uavs focused on quadrotors.

When Maximum Entropy Misleads Policy Optimization A survey of path following control strategies for uavs focused on quadrotors

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.818218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.533147Z digest=sha256:1136e0686010c75fbee67b52793c8fb55853e3bf3e031807a035fd66eb2c472c

Observation a98d043c-5884-4544-9165-a23de4ab58f2 · outbound

This paper cites Trust Region Policy Optimization.

When Maximum Entropy Misleads Policy Optimization Trust Region Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.535372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.535372Z digest=sha256:9e2236f689dabc05ca38499779949b925d761002dc1488399fafaf8179905b4a

Observation 68ca8698-a3ce-4e2d-a3de-af39c3fbb52f · outbound

This paper cites Proximal Policy Optimization Algorithms.

When Maximum Entropy Misleads Policy Optimization Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.537569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.537569Z digest=sha256:1409d40a0b9c85ca23226c7dcf93f615eb6c079072dc2627bee82c0e7a504db0

Observation 4759333a-bd67-41eb-9a2e-1ae8992324b4 · outbound

This paper cites M., Vergara, P.

When Maximum Entropy Misleads Policy Optimization M., Vergara, P

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.810564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.541152Z digest=sha256:744aa3af1ec716edc50df6c6e63cfc4377f85f8c5365fa80043b90d778b6e63b

Observation b1ee5d6a-f726-4b09-be7f-bb7e24cc38b9 · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:19.802698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.543450Z digest=sha256:2ee174c6cced8fafd58134545e43af0171338b351689648b13f9062601621151

Observation e5f9703d-9a98-4688-ab6b-7cdd8c8ac3fb · outbound

This paper cites Is robustness the cost of accuracy? – a comprehensive study on the robustness of 18 deep image classification models.

When Maximum Entropy Misleads Policy Optimization Is robustness the cost of accuracy? – a comprehensive study on the robustness of 18 deep image classification models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.794990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.545689Z digest=sha256:4f79558483c3ad52cf9e8b0fd1fbc7060c6f82705cff8c38aea96cf540616aeb

Observation 9f259164-14cb-49b4-a806-ad92c4130356 · outbound

This paper cites and Karak \"o se, M.

When Maximum Entropy Misleads Policy Optimization and Karak \"o se, M

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.785849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.547605Z digest=sha256:76dab2965b3055ac6f895bc8e9d63a7b27be0627c5e932d24091f943dbb080a4

Observation fec21014-f7c5-4cd7-8b51-d30639b59396 · outbound

This paper cites Mujoco: A physics engine for model-based control.

When Maximum Entropy Misleads Policy Optimization Mujoco: A physics engine for model-based control

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.549765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.549765Z digest=sha256:6b051e572dab3a5af218bc2458a7a5f23b7c0d99e88f89388667eb3f3869c249

Observation 278d8bac-202a-4fff-950e-61e269332467 · outbound

This paper cites Robot trajectory optimization using approximate inference.

When Maximum Entropy Misleads Policy Optimization Robot trajectory optimization using approximate inference

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.773741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.552457Z digest=sha256:89394388f0a9b5f1eccd54d06b5e05133765f478c32e4923a4acf3bb7c596595

Observation fc58c466-bcb2-4fb0-bf7e-757093231711 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

When Maximum Entropy Misleads Policy Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.555066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.555066Z digest=sha256:6d6cebb5c8cbe10ca127222a4bdb6bf3c3aa632c962e674f77f223e97b227fae

Observation 446f0514-18ab-48ba-84cb-5b9db5ec58ec · outbound

This paper cites Robustness May Be at Odds with Accuracy.

When Maximum Entropy Misleads Policy Optimization Robustness May Be at Odds with Accuracy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.557477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.557477Z digest=sha256:289415e1073adfd2e8c607e34ac3f6ba97cae99fc6d0c4846aa812db01b3ac1c

Observation 2c78c91b-c85e-4c26-90a5-9721dbf43dd2 · outbound

This paper cites Meta-SAC: Auto-tune the Entropy Temperature of Soft Actor-Critic via Metagradient.

When Maximum Entropy Misleads Policy Optimization Meta-SAC: Auto-tune the Entropy Temperature of Soft Actor-Critic via Metagradient

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.560314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.560314Z digest=sha256:b407ffabf2ef60895958e805e34c2d1206c0942d0ef9fd3abba24d61ef6398e0

Observation f21f248e-f555-4e3d-a345-a0a431a0a37a · outbound

This paper cites Tianshou: a Highly Modularized Deep Reinforcement Learning Library.

When Maximum Entropy Misleads Policy Optimization Tianshou: a Highly Modularized Deep Reinforcement Learning Library

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.562842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.562842Z digest=sha256:14a2c587b0559c27e9a8fd2725f21532eacd49e17c56be3978a3a10efeacb73d

Observation a40286e8-707f-43ee-a7bc-6c6d11661dc0 · outbound

This paper cites Karting racing: A revisit to ppo and sac algorithm.

When Maximum Entropy Misleads Policy Optimization Karting racing: A revisit to ppo and sac algorithm

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.765245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.565395Z digest=sha256:fd7836a9818844badc5c73c8184e1c4ffe6b1619f6ad7bc19f0c0f4614e6656d

Observation bd862ccf-0fa6-4ba0-9a40-62e9f5ef98de · outbound

This paper cites A Closer Look at Accuracy vs. Robustness.

When Maximum Entropy Misleads Policy Optimization A Closer Look at Accuracy vs. Robustness

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.567927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.567927Z digest=sha256:03e623172a6dd1020035a932a9901473c90f5ebc333efe325924de38fc3f2fe7

Observation dc757881-c934-4bf9-b07c-fa46c66f7b6f · outbound

This paper cites Theoretically principled trade-off between robustness and accuracy.

When Maximum Entropy Misleads Policy Optimization Theoretically principled trade-off between robustness and accuracy

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.757149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.570121Z digest=sha256:14f7f01ba0cb66fb3c65aac3b363c8ab74e6ebb3198d36657f8ceee79edc78c4

Observation 376d0516-78de-43d3-a668-855b9b69e303 · outbound

This paper cites Humanoid Parkour Learning.

When Maximum Entropy Misleads Policy Optimization Humanoid Parkour Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.572182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.572182Z digest=sha256:174fd5fae33f8605ba8d729aec3174d393fa946fa532c31f0f59950633c6b313

Observation d4fbaa72-699b-40c1-ade6-161f429e7168 · outbound

This paper cites an unresolved cited work.

When Maximum Entropy Misleads Policy Optimization Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:19.574355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:19.574355Z digest=sha256:eeef107c53283ec7a40ea03a3c1e57ccd588a4c16458deaa2dbc9b4d005c7225

Observation 501bdbb6-3a21-4c66-8c99-9df64c8a0f33 · outbound

This paper cites D., Maas, A.

When Maximum Entropy Misleads Policy Optimization D., Maas, A

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.745549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.576834Z digest=sha256:141b66a0a7b30426199ee88c7167156b58d92e4e9594edbe2f88177765d3df33

Observation 56f51a5b-be0c-4dca-992f-1a07088813b8 · outbound

This paper cites D., Bagnell, D., and Dey, A.

When Maximum Entropy Misleads Policy Optimization D., Bagnell, D., and Dey, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:22:19.736588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T10:22:19.579220Z digest=sha256:b52a844281e7ea155a972aacaed4ada8acd8d382ab3ab74fb7879d2597e74538

Pith citing papers

Observation 0e5ebc49-9234-42de-9a5e-aea75042f0fd · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF When Maximum Entropy Misleads Policy Optimization

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.807781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:fbd864ae35279a7a43845f2b05eade0135a643ee3c7d4a3246c8bd5a1c387207

Observation f247a322-b1c6-4be2-9a99-b97a6a08b6c7 · inbound

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization cites this paper.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization When Maximum Entropy Misleads Policy Optimization

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.858360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:afa0bea1d7dbe5d850d31c17809ebb1efddceb1c769177d878cd9b89bda4bc7e