Pith. sign in

Paper Citation Record · LEDGER

Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2211.15144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.15144 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:57:48.440353Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:15:32.316103Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5e4bbf9d-5b3a-4123-8bbd-95e07e88f81b · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 258

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:15:18.597937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:b1579fbb0969cc5b64865dde11f06f0361fa7e271e5ce7670c6eee262d618694

Observation 58ca045a-41a7-448c-824e-fd19dc28e8d8 · inbound

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents cites this paper.

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:48.440353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:57:48.440353Z digest=sha256:fad02cdd2ac700a3be86940a3d3e5c879e49109f7d401c805504a9ffc7c8bf3b

Observation d83cb6d7-1eef-42bd-996c-db7aca6e49d4 · inbound

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control cites this paper.

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:07.778134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:07.778134Z digest=sha256:3617bc7c80d73586c37b3081cf562f03f87b5e5010f2455e16e39a92cec1fb29

Observation 58facab5-fe10-48e6-b4a2-c0a0c40ca365 · inbound

Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning cites this paper.

Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:21:13.579804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T20:19:28.727598Z digest=sha256:30c32834ba7783d133cdd523c4677bdd75d9961d4c622fb1a251718765c2bc56

Observation 547a2e29-6d8e-4bfc-842b-cd9210aa0923 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:12.509676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T19:31:38.069592Z digest=sha256:4756d8db06402ab6da8f22fa34723682641c49176fc2d3c29e8167fd149a5c22

Observation 2c291311-8889-4673-b059-44ecf5f1b6c6 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.319056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d0f63b149fc7ab516a2b0c2a3c2e3e2d22ae165cfe58214a48efe9b877d02eec

Observation f633221a-fd9a-407d-8c77-5a26d9c64a93 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.648608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:ae91b1bd10a5f4bd92b6a5bb67da8c001fcb6103d7f950a101b5f28fd47c4238

Observation 43bb3947-5eac-42dd-8b84-b09c3f92b383 · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:58.205675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:d19ff7af82e7e28c2c9818fb016ca2fce5ea41ce30b241d8494bfb4d4997e53c

Observation b8fb2103-dbef-4a24-bf4b-722f0b0475d3 · inbound

ReBRAC-v2: The Return of the King cites this paper.

ReBRAC-v2: The Return of the King Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:46.715942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:32:46.715942Z digest=sha256:6494c0df2818d079a1a2e6d44628d94ba7914035d3f1f4ea75896eb23bf48945