Pith. sign in

Paper Citation Record · LEDGER

Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2211.15144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.15144 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:48:46.826872Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:15:32.316103Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5e4bbf9d-5b3a-4123-8bbd-95e07e88f81b · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 258

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:15:18.597937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:c6f39b1a39a8cf9069c975ac9b9684e135ecdf7b97ba6c577f21a2945355cdc8

Observation 58ca045a-41a7-448c-824e-fd19dc28e8d8 · inbound

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents cites this paper.

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:48.440353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:57:48.440353Z digest=sha256:eb77bef5a5a69da9bca2b09161644aebdacb861e364176f48ab925739d39637b

Observation d83cb6d7-1eef-42bd-996c-db7aca6e49d4 · inbound

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control cites this paper.

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:07.778134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:07.778134Z digest=sha256:6cbb60f10369f3223f7e5288c3f71c6126a8371a65d2e053ae6b1c98b434cb44

Observation 58facab5-fe10-48e6-b4a2-c0a0c40ca365 · inbound

Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning cites this paper.

Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:21:13.579804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T20:19:28.727598Z digest=sha256:7e9b73361e0dd7feb9026c46924ccc9af86d9821bc7eae2c1fd911d66cddf712

Observation 547a2e29-6d8e-4bfc-842b-cd9210aa0923 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:12.509676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T19:31:38.069592Z digest=sha256:2bfa012c743c87f4aa5d0d38735ef4be5a17dd11bfaf5f2ca95986d1de96d23d

Observation 2c291311-8889-4673-b059-44ecf5f1b6c6 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.319056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:42229ee500e55f34a79149c1cddd532f1542853d4ef9e1f4eb2baa4d534473ad

Observation f633221a-fd9a-407d-8c77-5a26d9c64a93 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.648608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:c4854a8af9cd6d5f925ecbdd1c3ee3c7c69723f9bbab1e6f7366206492810e51

Observation 43bb3947-5eac-42dd-8b84-b09c3f92b383 · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:58.205675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:2a1e9a75e9fd9de7c1fa564035a811d0f4a4a6c31b762f397896b05bc7bc052d

Observation b8fb2103-dbef-4a24-bf4b-722f0b0475d3 · inbound

ReBRAC-v2: The Return of the King cites this paper.

ReBRAC-v2: The Return of the King Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:46.715942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:32:46.715942Z digest=sha256:2bd857894e7590476d1794307edf020ffb2c2eb0527f7139dcca9978236c9a7d

Observation e6323f33-756c-425e-960a-7c7f7e0c19df · inbound

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control cites this paper.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 261

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:46.826872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:46.826872Z digest=sha256:006b17d31d64f97b751bbceeea2850f2e8db0d304a25ee5fdf7b8fddead615b3