Pith. sign in

Paper Citation Record · LEDGER

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network

As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2502.00288.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00288 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:39:58.748802Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact5
  • verified fuzzy24
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 91374a6e-fc4f-4e94-bfff-27c6b7243ffc · outbound

This paper cites Layer Normalization.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.569489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.569489Z digest=sha256:ad74490f63d92d4833b542f63c2179f4e5856689f05e6e2967d544f2e2984a18

Observation ef8aae6c-ceae-4080-bf79-253b2d321ea1 · outbound

This paper cites J., Smith, L., Kostrikov, I., and Levine, S.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network J., Smith, L., Kostrikov, I., and Levine, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.974310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.574141Z digest=sha256:80c8b41d94d2e7caadae1697d0fa40836ac600f1056e92ebc97a52522da8ab7b

Observation af8f9d16-85ec-47ec-8e05-7cc9d5413d18 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Dota 2 with Large Scale Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.577776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.577776Z digest=sha256:7b76f6c7b95a620ff585ad98b7750d9a6d72b9eee5f4780e6b18af089f246b9a

Observation b2f8dd92-cde7-4946-9ec5-23b8f5711018 · outbound

This paper cites Offline rl without off-policy evaluation.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline rl without off-policy evaluation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.963751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.582381Z digest=sha256:e429eeae144c9f2b6be77214a8c2f5bf11a72ae5f808554d39a20532a0b4be54

Observation a4fae73a-bb69-4f9e-a2e2-7d7a2fae07cf · outbound

This paper cites an unresolved cited work.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.586110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.586110Z digest=sha256:836680b2513601c9e2416b9aa7896f0b7c085e20aed69927c4cfe728097e0fd6

Observation 0150fed3-b684-4b76-8ff9-00776e071b32 · outbound

This paper cites A., Salazar, G., Tran, H.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A., Salazar, G., Tran, H

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.945981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.589949Z digest=sha256:e5021d37bcf740f6dcd98e1970db766006082ffaa48ec9e540d6f074dbcb5f66

Observation a4c28b55-7c16-4d0e-bbf5-c61969531f0d · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Decision transformer: Reinforcement learning via sequence modeling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.935274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.593890Z digest=sha256:36d5fd2ba59ec75187bba35bdd6f6d42ee8accc3084a44f74c374af476e1c5bb

Observation b0d109ad-e96a-4b00-994e-d6727e0c5b54 · outbound

This paper cites Rvs: What is essential for offline RL via supervised learning? In International Conference on Learning Representations, 2022.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Rvs: What is essential for offline RL via supervised learning? In International Conference on Learning Representations, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.924194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.597241Z digest=sha256:6cbddef7215c1451cf2de6fc6ed86dc3df80c5eac3c893ffeecf1cbdb5b2659a

Observation 1fecd2b1-1114-40f1-a0e2-6ef574376969 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Counterfactual multi-agent policy gradients

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.600622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.600622Z digest=sha256:b3dc6ceb0771cbb6a8ed441779191405e60a64cfcc7a54ecaa10de7c83ca166d

Observation db3b8a5c-55dd-40ad-a00f-948b388fe3ef · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.604186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.604186Z digest=sha256:cf6f48ac7e40e043ba261e6d2db0fe331edf27645a3b58df0fcea15caf8cd159

Observation 7cf1abf2-b44b-437b-90f8-0fe4bb516db4 · outbound

This paper cites and Gu, S.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Gu, S

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.913496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.608251Z digest=sha256:2fc060b779e8efd173f57a82eda966fb595901bf29b4689a34eed686fbaf09dc

Observation 2e974eba-5512-4ef2-9190-856862d12ee9 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Addressing function approximation error in actor-critic methods

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.612097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.612097Z digest=sha256:5074bc7b14df5d6e5a5358554aa35b8e42a9e6cd0a72b0390b20f7fb32f9665f

Observation 26ad1416-bb4e-4828-8072-f452aba5f1ce · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Reinforcement learning with deep energy-based policies

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.615756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.615756Z digest=sha256:7ae39ea515f932db7723a050bc381a25921c5429ac1fe3b5eb1ededa8d308f99

Observation 2e7cc4ae-b972-48a4-be71-8365288c2187 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.619235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.619235Z digest=sha256:c9c75173a93367693b3648f160c9b952a967c72467c128905373576f3451d29d

Observation 762c91c2-2744-4237-b726-5ae2ce80c82c · outbound

This paper cites Modem: Accelerating visual model-based reinforcement learning with demonstrations.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Modem: Accelerating visual model-based reinforcement learning with demonstrations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.882950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.622632Z digest=sha256:36d8f15364c08e1870c41fc099dbc8c9f8fd1c143173025f49dc8daa8925df3c

Observation bd92b01e-1c3a-4daf-b9b1-0e7201050ce0 · outbound

This paper cites Neural networks: a comprehensive foundation.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Neural networks: a comprehensive foundation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.871795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.626098Z digest=sha256:cc840bc17ede1a9f723149c082b705c1800b91513780a80316f485941edbfd1d

Observation 12eb4e6e-2e22-44b3-ad2c-b7be9411f06e · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Gaussian Error Linear Units (GELUs)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.629593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.629593Z digest=sha256:874185fcfa05a3f2bebafb63e840cf33db99a7ab23f118a79d925106c52484b9

Observation 7a2b416e-2a8f-4220-b613-e2dec9408142 · outbound

This paper cites Deep q-learning from demonstrations.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Deep q-learning from demonstrations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.633495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.633495Z digest=sha256:5cb2603a40db51620f991a798e054f87adb7c5c0390039d4cf94f87c429c8a55

Observation 6c80aa3c-ffd2-4295-be7a-5301295dc6a5 · outbound

This paper cites B ayesian design principles for offline-to-online reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network B ayesian design principles for offline-to-online reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.860365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.637064Z digest=sha256:fee6d7811c3c9dfeb854f1431f738d7831084c870bebfa75442de56790bac922

Observation 493f95b0-9b97-4569-bfea-0995ca08aa92 · outbound

This paper cites R., and Davison, A.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network R., and Davison, A

Reference 20

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T19:39:59.593966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.642052Z digest=sha256:2ce866059e839d391addd85d37cbd18d459c8219664d3fce390e31a0315e4e9d

Observation b47e30cf-10fb-42c2-b51d-1a33e754ab05 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline reinforcement learning with implicit q-learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.645779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.645779Z digest=sha256:c17a23f8d83c01d5ed785b461d13c34d9bfc1363ceba853fb575d501e44afcbf

Observation ccf644af-9f28-4a5f-b182-1954113cb865 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Conservative q-learning for offline reinforcement learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.649212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.649212Z digest=sha256:4cfd87751ccd817457422eb1667cbac30553d0d2d2eabd5b3ef506e9bf16abe9

Observation 538adca6-c19f-489d-9158-65b072b067f2 · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.837241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.652437Z digest=sha256:8a84f1b9cc35359b70d15ee922071b999338329b6faeb0270d55ef8be6764929

Observation aaf53e1a-99d2-490a-aa35-c56a737c90fb · outbound

This paper cites Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.826694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.655705Z digest=sha256:9d00eb9f21329c3cfcfb6ca2f81677ef28f5b73353d1760a102b05c5f1e4badb

Observation aaf30073-6ddb-40f9-9145-f0178adef183 · outbound

This paper cites A survey of convolutional neural networks: Analysis, applications, and prospects.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A survey of convolutional neural networks: Analysis, applications, and prospects

Reference 25

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T19:39:59.411897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.658938Z digest=sha256:676dfcddafeeedee2e9b013253ef74d1e9a3e8630e8fc25877230bdb6499ea6e

Observation e884384c-00bc-4e93-aeb3-71aef5333199 · outbound

This paper cites Continuous control with deep reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Continuous control with deep reinforcement learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.662257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.662257Z digest=sha256:ad9c647efe678dd1a86023413d6b091c10ff8a847566534c735fcee4d3d34a45

Observation 824cb68b-7c81-4977-99f0-f85c9e5c024a · outbound

This paper cites Discrete Sequential Prediction of Continuous Actions for Deep RL.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Discrete Sequential Prediction of Continuous Actions for Deep RL

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.665938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.665938Z digest=sha256:9c02d5724de16ffcda2b444ef83ed9fd82840db6d0053302eedffe463c8f9e6b

Observation d4425530-19bc-4373-a556-f48c66a2efe0 · outbound

This paper cites A., Veness, J., Bellemare, M.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A., Veness, J., Bellemare, M

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.669816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.669816Z digest=sha256:17995ff24aff91eee14514dfbfe62222d754994b955834024ff6eeacfd28ba6c

Observation 02cadaba-4b3f-4c95-bc9b-b71a628de3d4 · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Overcoming exploration in reinforcement learning with demonstrations

Reference 29

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T19:39:59.176348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.673153Z digest=sha256:f5ae8c23e666ce61c05d8187dadeb6a0ec184b6e087c237346ba4c2702f84730

Observation f3d12f85-0098-4afe-b502-e03eb612d060 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.809768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.676302Z digest=sha256:a4e61adea43bcd9170f3e82d6e0e05595f944cacc1fc7b7088ab41f46dd34d55

Observation 52cea845-30ab-41f4-a614-e4c62b127473 · outbound

This paper cites Learning complex dexterous manipulation with deep reinforcement learning and demonstrations, 2018.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Learning complex dexterous manipulation with deep reinforcement learning and demonstrations, 2018

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.798314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.679741Z digest=sha256:b3a23219981065ae89ae9ba7758fed51fe2d064d52bdc814a335fb9dc84d8f31

Observation 991c977d-4f6b-4758-b19a-9f829615dd29 · outbound

This paper cites an unresolved cited work.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:39:59.787378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.682953Z digest=sha256:1c47b4e2fb590edf5ec19a0ee4a8d0589af955a029f9191df739ab2c39daed98

Observation 358af754-ac01-460a-8232-e75feb34d301 · outbound

This paper cites Mastering atari, go, chess and shogi by planning with a learned model.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering atari, go, chess and shogi by planning with a learned model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.686288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.686288Z digest=sha256:f459897b6ca9ea6eb36e0a9090f32dd435024d0261d8856051b241ca352b4ed4

Observation ceb203cc-861f-4e3c-80ee-a5fddc944378 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.689604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.689604Z digest=sha256:b33070900c9dcdb78132d1b7b33ceed4051badce56b1501f6c5667d63a2e3d76

Observation b3b41d11-1f2b-412f-bcda-c5f7599ee9d2 · outbound

This paper cites and Abbeel, P.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Abbeel, P

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.693154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.693154Z digest=sha256:3a50b54c04b04e083d77d805595e1620e39520fc719b40047d41602cb9cd4512

Observation c10b52f3-7744-4a07-94d5-28a36404e590 · outbound

This paper cites Continuous control with coarse-to-fine reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Continuous control with coarse-to-fine reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.769503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.696684Z digest=sha256:5e15db7bcc8ac08d0f8b0692644b62a0ed0cec1fd25a657167c58e2a3a9e9e4d

Observation 26766bd2-613b-4f3f-bc8a-b5711008786f · outbound

This paper cites Solving continuous control via q-learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Solving continuous control via q-learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.758043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.700226Z digest=sha256:50e8b581fe99cff111a3f4b54ccd4e9a30623b8766c6fefa80d58e3b4f3e6d94

Observation 04f49af1-4824-4fed-b2dd-03e662db1c07 · outbound

This paper cites Growing Q -networks: S olving continuous control tasks with adaptive control resolution.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Growing Q -networks: S olving continuous control tasks with adaptive control resolution

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.746815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.703922Z digest=sha256:1e71994a645bd5f02168b29abf5ab1a93a7a47a9dc717525442e6e20ae6b0031

Observation d5ef7617-735e-416b-94eb-1fa96bc507a2 · outbound

This paper cites Mastering the game of go without human knowledge.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering the game of go without human knowledge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.707331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.707331Z digest=sha256:85c9463feac047cb454757590f5803695817323a1bc6372a5d7275bf680bfcf4

Observation 590cead1-f946-4069-816b-55fde7943841 · outbound

This paper cites and Agrawal, S.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Agrawal, S

Reference 40

Resolution
verified exact
doi, observed 2026-08-09T19:39:58.802667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.710807Z digest=sha256:c005d60916bfdab9d2e898de3f28fb5ce95ea3cbd49af3680ea8641a499ac3c2

Observation b2aca0e3-7fa3-4625-9fa3-72622036c302 · outbound

This paper cites Action branching architectures for deep reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Action branching architectures for deep reinforcement learning

Reference 41

Resolution
verified exact
doi, observed 2026-08-09T19:39:58.790610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.714256Z digest=sha256:d9a35c51d910aeb7ad22e34c7309d41c77bee7674153484f594bf3428d5590d6

Observation 76bb7886-7220-4fcb-b150-e0503c4d87ab · outbound

This paper cites Learning to represent action values as a hypergraph on the action vertices.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Learning to represent action values as a hypergraph on the action vertices

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.729926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.717906Z digest=sha256:74ae19f8f50424d98abbcafa319ec6dbbf4cf37dd5aa06ca9adab74e9934130e

Observation c3230589-da4a-48a3-a5e1-f6b7a96708c4 · outbound

This paper cites Deep reinforcement learning with double q-learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Deep reinforcement learning with double q-learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.721491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.721491Z digest=sha256:eaad0ed948720a451c590fbfd2abdf1384c0356bf592933c7d4d85d44d6fcc1b

Observation 909c3e06-0f3d-4fa0-a2cd-0c5d93d6deb8 · outbound

This paper cites Discriminator-weighted offline imitation learning from suboptimal demonstrations.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Discriminator-weighted offline imitation learning from suboptimal demonstrations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.719141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.725163Z digest=sha256:695e7992f630e25d69a4780d36170b208b3ffdc4bff4c4105e0ee157e3793485

Observation f2673bbe-0e28-409c-acef-3c4fb7a9d297 · outbound

This paper cites Hd-cnn: Hierarchical deep convolutional neural networks for large scale visual recognition.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Hd-cnn: Hierarchical deep convolutional neural networks for large scale visual recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.708249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.728435Z digest=sha256:7116189c1c0e79a047ea0df8dc0cb771f30370c7c4eb905f7787bcbe1f419942

Observation b7835a39-4772-4a97-9a2c-832d712639db · outbound

This paper cites Mastering visual continuous control: Improved data-augmented reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering visual continuous control: Improved data-augmented reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.695917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.731948Z digest=sha256:16ea4102f2e15c6b9929a5517f9132a8cdec5d45f46fe57d8eaefe8a342c2055

Observation 51b5b9c5-0637-44b1-8770-9a8c6aa53510 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network The surprising effectiveness of ppo in cooperative multi-agent games

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.684908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.735090Z digest=sha256:79d322644a99cbe10a4bd70b76b04da6a3764ec65c3860e9ccc123f4fc850a2b

Observation 0094e1ab-6b8c-48c7-92db-8cf0d2792c8b · outbound

This paper cites Policy expansion for bridging offline-to-online reinforcement learning.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Policy expansion for bridging offline-to-online reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.674731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.738490Z digest=sha256:11feeb2e3889f9dcd54d828c26e5eff8272b92d60fa84d7acc0518b3f03a3751

Observation 87a41440-f206-42de-8260-e1bae4ec3c3c · outbound

This paper cites Z., Kumar, V., Levine, S., and Finn, C.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Z., Kumar, V., Levine, S., and Finn, C

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.664137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.741871Z digest=sha256:6e11e053be9853b69b03831c85ba19d544e75df2db8d1887200a3304c7bc068a

Observation c7add5a0-d9ee-428d-8d03-a620df032e05 · outbound

This paper cites D., Maas, A., Bagnell, J.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network D., Maas, A., Bagnell, J

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:39:59.653624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T19:39:58.745240Z digest=sha256:f0c0557c825899adfa526e9a2eb3ca9661a07ad49f462e8f329a86126100b2bc

Observation 4a2aa577-c749-4a7a-bd3d-e5e44f236f26 · outbound

This paper cites write newline.

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T19:39:58.748802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:39:58.748802Z digest=sha256:43ef1fa34ac494ce7658ed5cd002f3007dc1fdb1c70ce9a76fdd97733d954bb1

Pith citing papers

No inbound Pith citation observations are available.