Pith. sign in

Paper Citation Record · LEDGER

Extreme Q-Learning: MaxEnt RL without Entropy

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2301.02328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.02328 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:38:30.228255Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.943655Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 575ff729-ed03-4652-b4e6-6f605c574e73 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Extreme Q-Learning: MaxEnt RL without Entropy

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:48:36.476653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:069248fa9722145c87f1681c4c8adc0a83d38ca5afa46782bebae8151245c114

Observation 8b590690-0025-41e2-9905-07b1a5960926 · inbound

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint cites this paper.

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint Extreme Q-Learning: MaxEnt RL without Entropy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T04:38:30.228255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:38:30.228255Z digest=sha256:782664908ee40f749a12de796392ca1994ce2ddc84bd282e7ee3cc3160631d50

Observation 202fe6ac-af03-4756-8236-36e4ef3b38c3 · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.632354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.632354Z digest=sha256:11fce1bd2f8f5a1aca847f37218fd9212fb8ec513a2a6b5cd7f9f706fdcd7c73

Observation 14dd046c-95d9-4c78-9362-9a3b7e400315 · inbound

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood cites this paper.

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood Extreme Q-Learning: MaxEnt RL without Entropy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:24.560434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:22:24.560434Z digest=sha256:2fcbfa51d9d28df3c208a33b133e98630c73f8e981bbc845d32f3ea4178b0426

Observation 6b556c63-0a75-4c9d-a391-14da58e48501 · inbound

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion cites this paper.

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion Extreme Q-Learning: MaxEnt RL without Entropy

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:30.548506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:30.548506Z digest=sha256:5d7e79e8589cc49b5f888b07096b446225533149adcfc59cd17cb27ecd6fcf9e

Observation 09ecb643-05b9-49bf-94ae-8708d9a5997a · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Extreme Q-Learning: MaxEnt RL without Entropy

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.587307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:dd2da74c283e4e8f95bcb05e8f679a61afdffcc3418b2feffc7cac5457f2f291

Observation d11e5360-8f6a-457d-9fe9-64a546f422c7 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.657389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:0dd7ef1d05207a8c5811554939644bf8365cb5d5b19864fb8de2204f3ea47f4b

Observation bf7a61dd-b81f-4142-a925-aa35fc35e55f · inbound

Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer cites this paper.

Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer Extreme Q-Learning: MaxEnt RL without Entropy

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:10.148459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T13:56:01.365382Z digest=sha256:f82be5a880b7176eaefa3a659dd152a9dbcf671d67bafb6e404a483af5584fce

Observation 3d758c3f-83aa-4a98-86e0-7347ec0464ca · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.853436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:6200b41d8fbeac6c2d8d0063baefec5337fdfa01e630cec101d32bfa069a05f7

Observation 0729d11e-2144-4596-9322-bf9bc9c14bbb · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment Extreme Q-Learning: MaxEnt RL without Entropy

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.924661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:2d8cdffc2b685e4d48f965b254051598280d92c3d78073e7d2f9b9d73b4bc1a4

Observation e8b078c2-2a53-450e-8180-6e537c3574f3 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.945140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:ff66ec4964d773c79de6a1588f6009c5a5143a625039da4f87d50c4bae461313

Observation b8ac3f31-467e-42f2-b8c9-b082063529b4 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Extreme Q-Learning: MaxEnt RL without Entropy

Reference 269

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:22.550304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:22.550304Z digest=sha256:20d08cc7cceb88a76006cfcc3eefcbab0c29255a104ebcf103e62ed03e64667a