Pith. sign in

Paper Citation Record · LEDGER

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

As of 7 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 9 inbound Pith citation observations for arXiv:2505.23150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23150 v1

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:54.542081Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:05.516882Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T23:09:12.737900Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved76
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c403830a-7572-46f5-8a4c-4e8f99131eac · outbound

This paper cites GPT-4 Technical Report.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:44.854747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:44.854747Z digest=sha256:c186946f15c870ed0570cd753d18f8371e5e90f94dfed3f9f7ad3234edeeba38

Observation 80c9da86-63a7-4a64-b95f-3a0d48e2c8c8 · outbound

This paper cites Provable benefits of representational transfer in reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Provable benefits of representational transfer in reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:44.967526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:44.967526Z digest=sha256:777163462af83ad9492aa043906b59523b9d3737d6018ea35c0d3e67cc105294

Observation 713717d6-e6bd-4c9d-96bd-c26975730a9d · outbound

This paper cites S., Courville, A., and Bellemare, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Courville, A., and Bellemare, M

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.019350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.019350Z digest=sha256:12341767f92d165ab2ffcd5e5e2a288aa99fa6107f21bcdeb4711daed786c6af

Observation 97a075bd-184a-4409-aee0-3305870ecba4 · outbound

This paper cites S., Courville, A.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Courville, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.121469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.121469Z digest=sha256:010f582d79fa93c8176e46404f55a06258d9584dc933bce6a153763fb763a94e

Observation 7da3fa4c-e85c-4342-99ec-4195f8b9d31f · outbound

This paper cites Hindsight experience replay.Advances in neural information processing systems, 30, 2017.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Hindsight experience replay.Advances in neural information processing systems, 30, 2017

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.197775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.197775Z digest=sha256:2cb49c2ca16f34cf7435b8442c313cec310659c0f7c30ec9690a233c9062986a

Observation bf984bb8-9e57-4b9d-b244-bc6220398bcc · outbound

This paper cites M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.269538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.269538Z digest=sha256:5f94c30a77e42e3bd055370279d83db22d04f1947dc080591abc28254990283e

Observation 74c0d758-659e-43e9-ab11-00f55115d6d5 · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.345247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.345247Z digest=sha256:5ef22c01d281cd0bccbcbc797d5832686c44abed1ff211f534cb65d4cf399f51

Observation 98169282-9aa4-4dae-8152-7ca68d2c59fc · outbound

This paper cites J., Smith, L., Kostrikov, I., and Levine, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Smith, L., Kostrikov, I., and Levine, S

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.433639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.433639Z digest=sha256:e6a91382b0f2ee94b0a4d95501f96265f9fb6382d4dc01101fe09459a601ea98

Observation 3b1b6998-d6db-4859-ba31-625e5dc47236 · outbound

This paper cites J., Schaul, T., van Hasselt, H.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Schaul, T., van Hasselt, H

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.571178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.571178Z digest=sha256:eb887920b8901bdb564e2ae57a54d697f6e22ec8bb7cd32fe7b4d7d6423afb75

Observation d9949fcb-109c-4545-853c-02164a0e7727 · outbound

This paper cites G., Naddaf, Y ., Veness, J., and Bowling, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners G., Naddaf, Y ., Veness, J., and Bowling, M

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.663223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.663223Z digest=sha256:dd0bf2f32db04fff2d38914884c59eec78bc9d846a54218730483fbee6697b80

Observation c5f64815-4508-4e5a-ac03-e3cbf631d284 · outbound

This paper cites G., Dabney, W., and Munos, R.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners G., Dabney, W., and Munos, R

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.750044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.750044Z digest=sha256:9ceaae20feda61a910a2df4bf5bbd3774e7df95c29d7def538e753b4655886c3

Observation 08d1053d-46cb-4f71-880d-78bc3dd62be7 · outbound

This paper cites Dynamic Programming.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Dynamic Programming

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.804154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.804154Z digest=sha256:e046a798dde3b1d663d7d409be19c24365d13a0c654fa000872a2e3f4418da1e

Observation 2486c5fb-c6b4-443c-8cdf-fd42dd633eb4 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.874874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.874874Z digest=sha256:74b06745239ae4205644cf7f75aad5a459a29a66a5d98002952a3fdbe6af555d

Observation 93b4fe42-9f73-4d76-82fa-1430bf502f95 · outbound

This paper cites P., and Weinberger, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners P., and Weinberger, K

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.968173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.968173Z digest=sha256:397f9f0669363fb049ad9d01e036c65458d90675b8461bb8e9dc363a41dbde68

Observation 68d41d36-8aae-423b-996d-11f1b1fdfef7 · outbound

This paper cites J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.065866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.065866Z digest=sha256:ac89793296cb6311f01bbf9e6674502583e306c81ec35f0ff18186d2c1a9ab49

Observation 63c3a67c-9aca-4c18-9056-0c246c26dc54 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners RT-1: Robotics Transformer for Real-World Control at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.158225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.158225Z digest=sha256:5c5dabd23dc981cdf31c8e753f3dc17bee27f13949ee8319ca1188e7b4af0bb3

Observation 778f8cbc-d862-4bf3-b45f-d40d66689849 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.274336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.274336Z digest=sha256:0807280b87618586774b6768c7174440195758993c1ba277e8a4d482d12b4570

Observation c8a57e69-4a22-4384-9865-e7bbadf3fcad · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Decision transformer: Reinforcement learning via sequence modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.379002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.379002Z digest=sha256:506d339af30d26bc3e4d0456c1cec409bb30a93a1deb31642a66041bda9a5c81

Observation 8bfdf1d4-ebbf-4fa2-969a-3b77060d52cf · outbound

This paper cites A system for general in-hand object re-orientation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A system for general in-hand object re-orientation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.445275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.445275Z digest=sha256:be46c00c4b5a408db9a04efdc39b36b671a0ce38ac7fbb90113a3215c3c340d3

Observation 5c71d3a6-7cbf-4467-85f6-f06c0eb2f829 · outbound

This paper cites Gradnorm: Gradient normaliza- tion for adaptive loss balancing in deep multitask networks.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Gradnorm: Gradient normaliza- tion for adaptive loss balancing in deep multitask networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.566082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.566082Z digest=sha256:8de9284e6c38b403202ad2a4172717c6d200e81bd773decd79c3953c2149aa9d

Observation b133e83c-fc66-447e-8f32-286d8c5f2621 · outbound

This paper cites Just pick a sign: Optimizing deep multitask models with gradient sign dropout.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Just pick a sign: Optimizing deep multitask models with gradient sign dropout

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.662012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.662012Z digest=sha256:9fe8a1ea9c3b9f1e80f995e13c1d7636fb72fc747fde81cb1a1b5c16b7d4d1ef

Observation 725d0859-c737-4840-9e5f-ba3f83d84306 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.707628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.707628Z digest=sha256:c4656672eb190fe41fec5411cc8253cd1d282a82e0f36b1f3f1f429b2175aaf2

Observation e5752454-cbfc-4efa-9bae-76e8fa6c8afe · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.750339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.750339Z digest=sha256:abd6722ae60c3ec3d845941627a9426267dc36ad1c86cfa9c813ad9b9accec4c

Observation f173b2ca-f1b5-45e1-a98b-0ac0ff439627 · outbound

This paper cites Better exploration with optimistic actor critic.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Better exploration with optimistic actor critic

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.835161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.835161Z digest=sha256:b6be53f7b969838de7eb51946caf9277979454ac9c5ef2ff759a85e288002161

Observation a9579dbe-bdc1-4982-8b39-ed18dca7ba0f · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.867259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.867259Z digest=sha256:1302d44d9c17db5922ceb3b23a20e310258f63bbd3fc18cabda2a6a61326cea0

Observation 544c6b33-a6c6-45fd-ac93-60cd1d1f1f1d · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.915785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.915785Z digest=sha256:c81c4b6eecfa653396c984592226c06125c3fc49f775febdcedf96ae8aea59dc

Observation cedc50c4-c18d-4625-9599-87af9b1a4611 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners An image is worth 16x16 words: Transformers for image recognition at scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.009204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.009204Z digest=sha256:82297809eaf31ae13b88846df63ed1e0537d6dd7c91fb9fbb225db32aff82ee6

Observation e88b3017-0087-4ffc-864a-353bd7bbd744 · outbound

This paper cites S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.060172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.060172Z digest=sha256:4de8677b882eac3bac546fa845acfce06f53764d15e1ac6be75207e85e8bf140

Observation 8156fac5-839e-4306-9452-eb846bba34f8 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.128470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.128470Z digest=sha256:9ce0370defaf9a29074655f888b755abe559786e77c70d6851b7e47d83d5dc88

Observation a990062a-6dba-4f00-b38e-c2f4518b63d4 · outbound

This paper cites A., Chebotar, Y ., Xiao, T., Irpan, A., Levine, S., Castro, P.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A., Chebotar, Y ., Xiao, T., Irpan, A., Levine, S., Castro, P

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.227015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.227015Z digest=sha256:06d1fa99a1d1afe9fad241f6e1df94d33f4a87e22fc5cb327c9e09693ac094e1

Observation ba3cede9-4995-41b8-a536-f8b58699f1f3 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Model-agnostic meta-learning for fast adaptation of deep networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.334044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.334044Z digest=sha256:bbefc3e6f3a27b9145294c23a30ec25b41c3460e41beb91917072ae88223f96c

Observation 5e78924f-1e27-4136-9ecb-d46590b3ef89 · outbound

This paper cites and Gu, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Gu, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.398149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.398149Z digest=sha256:41ac186090e3632c5bf311ddb202eebff854c2bfa86c7555a183e0b531e4a221

Observation b3f8af50-5032-4c35-8ef3-b0e167b5c22f · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Addressing function approximation error in actor-critic methods

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.450522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.450522Z digest=sha256:82601c138eecd8e7e4aa6ff955f008ccaeb474a14127ed421e7c293f59fbcd70

Observation ad6cab91-fa5f-41a7-bd2b-995743cf13f6 · outbound

This paper cites S., Precup, D., and Meger, D.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Precup, D., and Meger, D

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.514281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.514281Z digest=sha256:0744e6640c5007c5e012abda715f7a50e47c98f56e8ba25c508d1636e053ff89

Observation 3f5c7ad6-7ec2-46ff-b60d-2c762b03bdd6 · outbound

This paper cites Divide-and-conquer reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Divide-and-conquer reinforcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.599219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.599219Z digest=sha256:9cfaa78273d52f5da54d28510c34eaa361c84346eab1b83d8bd4b3e59682c41d

Observation 11695065-eb2b-43bf-8acd-b1b64740f6bb · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.702018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.702018Z digest=sha256:25c61b32fb19fba51e4147b22b2e2e22a7af039c3e8468b8dcb1e8d2c92ec0d9

Observation 9b11e6db-448c-4902-ad8b-699cd749dc7e · outbound

This paper cites Mastering Diverse Domains through World Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Mastering Diverse Domains through World Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.810738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.810738Z digest=sha256:829dc9e97205003e25c5ef223bfb1c55b43669900fa3769d63ca27edc495b45b

Observation dcd824f7-6396-4ee8-87f5-7728212e1b21 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.926005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.926005Z digest=sha256:f1d8323d1e068cf13de79a2644beb55c812dabe00060f131c98dbcb9de0a8260

Observation ee899869-f4ac-402a-bf1a-243769f12730 · outbound

This paper cites R., Millman, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners R., Millman, K

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.989117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.989117Z digest=sha256:f75b49b558915dbab71ac93798553c8f7ebe4398a3270bc60adb5add0e08bc1b

Observation 83679e9d-c067-4d62-96aa-c79d0a711e38 · outbound

This paper cites T., Wang, Z., Heess, N., and Riedmiller, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners T., Wang, Z., Heess, N., and Riedmiller, M

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.058887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.058887Z digest=sha256:a5608810a2ca4eaf2c86c21441466dd8436b3c68ef7da1ac94a0ccd03fda3b77

Observation 0e9bb63e-08e0-494c-bd15-61e1e3d3a57a · outbound

This paper cites Efficient multi-task reinforcement learning with cross-task policy guidance.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient multi-task reinforcement learning with cross-task policy guidance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.115400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.115400Z digest=sha256:e1ae32fa5e14fcaaf04392306d963f21f86dbcc79d92c3903b63844c462fe0e5

Observation b8e672c4-b654-4d62-9126-c99e3324d0bc · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling Laws for Autoregressive Generative Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.198228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.198228Z digest=sha256:1e6cfff704248141e8bdc1f86f10b62a11a17a78280961970431d6e5977f73f8

Observation aa26013b-9414-44dd-8134-fb1b1af1cdd2 · outbound

This paper cites Multi-task deep reinforcement learning with popart.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Multi-task deep reinforcement learning with popart

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.306053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.306053Z digest=sha256:d05bc6a8d3c21e210328eb73502efa949bf8bfe1284aa5c5194fe5621c08c460

Observation e10abf86-0652-4874-9896-dc073ebdf597 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Training Compute-Optimal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.423350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.423350Z digest=sha256:e2b6776539b34c92f97da5116c7b40e5a05b8f25a3959f558755a4d01bcd424d

Observation da2ff812-3cf0-4a19-a9f0-32166cf81636 · outbound

This paper cites Otter: A vision-language-action model with text-aware visual feature extraction.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Otter: A vision-language-action model with text-aware visual feature extraction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.535699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.535699Z digest=sha256:33096d0b6642e30232002fe56ad9e613de0442726b2bf536d0e388bbe2302e95

Observation 1d70b86f-53ea-46b7-86c0-9f050d402e4c · outbound

This paper cites Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.664444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.664444Z digest=sha256:f1cdb9612d08b5d1d2d8126e25b86733592016d5e70422f75bfa80dcdec30419

Observation 433abc24-6485-4cf0-9535-847bfb6ac080 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.794678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.794678Z digest=sha256:cbc0818e4fa7432bb3d96d9187906350e01547590ef4e51800bdd7f89eadf0db

Observation 1c78482a-fe9b-43f3-9780-a093bedcb51a · outbound

This paper cites H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.510973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:48.918826Z digest=sha256:2404d18099b36dbb310c21d460e57c81a6d40bf1456681b3e045bb08947484c5

Observation a618c6a6-b9fc-42b8-81af-bfefb5440eaa · outbound

This paper cites Scaling Laws for Neural Language Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling Laws for Neural Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.042717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.042717Z digest=sha256:f14be5e66674bea14fff4589fbd899c20bdbfbd70702f5646c194e2c80a0ca6c

Observation 5a2632db-b1fc-431f-b978-3e0be2c69394 · outbound

This paper cites Champion-level drone racing using deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Champion-level drone racing using deep reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.350502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:49.165271Z digest=sha256:3dda8d495a00b190dab2d3830263002ee67d88fae81baef65a5637de881e3889

Observation 9dd9a88f-9958-4e08-a061-06a00151d3e1 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners OpenVLA: An Open-Source Vision-Language-Action Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.306614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.306614Z digest=sha256:6086c42d02fca53328d0c25bd3493b55d7863a3f255faa35b91e2801a4ed164a

Observation 82ba696b-a068-4fdc-942f-828e6bd8b89a · outbound

This paper cites and Ba, J.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Ba, J

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.176243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:49.440203Z digest=sha256:c5b79e1e4241d281dc7d4e7abd2e927a7e79893d2942f6cb7b5eb4a34a41118b

Observation 25391644-5245-47ae-a423-bd5dfe19e699 · outbound

This paper cites C., Lo, W.-Y ., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners C., Lo, W.-Y ., et al

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.005942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:49.554910Z digest=sha256:8f914789a5f8f788bd840694b26e5b16fa1d275e30f97153c2d1885e46a3f1af

Observation dabe2b8f-9758-4483-88a5-a239cd795e92 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:05.816415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:49.704891Z digest=sha256:b83ac666b5124bf5f75723e2dafeda7ca9954105c677662b77b6647bbf64ec11

Observation 0b377d3d-6be3-4aa6-b19d-e7f2245a7783 · outbound

This paper cites Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.816226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.816226Z digest=sha256:2bbbf40db37ec8f66dbe07a9cfba1cee27b1d0830ec7eeb2a0fa4b28e1f040ba

Observation e0fce54b-56f3-4e46-9d78-a750c2e7257d · outbound

This paper cites DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.941083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.941083Z digest=sha256:561169c1c95c440832617bca1f49254249e4cfb16a9ec99b675443f267469ff0

Observation b92fdb5d-b34e-42b9-9dd8-0b79474c4011 · outbound

This paper cites RMA: Rapid Motor Adaptation for Legged Robots.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners RMA: Rapid Motor Adaptation for Legged Robots

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.021389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.021389Z digest=sha256:ac1d2a488458a732dc4770b250e3140107c72b003d040a5ff62f59405d0a468e

Observation b7b884ea-e9b5-4d08-8c9d-2d6a9dafe2a9 · outbound

This paper cites Offline q-learning on diverse multi-task data both scales and generalizes.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Offline q-learning on diverse multi-task data both scales and generalizes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.630424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.084622Z digest=sha256:bdcf55195ddf0cdc2935ee2ee4cc0a790f118547f6260d22935a20fa1e9aa53e

Observation e53f8be0-5c13-4972-8dd5-fe5ce4b48e33 · outbound

This paper cites Reinforcement learning with augmented data.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Reinforcement learning with augmented data

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.433016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.148267Z digest=sha256:021e97805b2913324d92ba9e9f1a210b2a2943340ff2a3105a8ed714e88994d1

Observation 6e598f6e-02ff-41e2-80f0-7de53dca1f64 · outbound

This paper cites SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.224254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.224254Z digest=sha256:8f6fbf3f5167d5ac61df899c14c9dc0440a511483c9e0e62c409f772667f30e4

Observation 0ac955de-f366-46dd-950e-f7f7e1d1f34f · outbound

This paper cites Hyperspherical Normalization for Scalable Deep Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Hyperspherical Normalization for Scalable Deep Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.269251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.269251Z digest=sha256:5003db772db431ee835da6a69bde99a9e2a95b46b2f7eba82f6241ae85c30952

Observation ae14b0df-74e1-4d93-bc6c-6e594c8ee1ac · outbound

This paper cites End-to-end training of deep visuomotor policies.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners End-to-end training of deep visuomotor policies

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.163798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.321875Z digest=sha256:34cafc140b05988ab8e51041e436f18657f8079d7297e8327666dd527bdaba89

Observation ba682875-3697-4628-a572-3f871e74de31 · outbound

This paper cites DeepSeek-V3 Technical Report.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DeepSeek-V3 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.379474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.379474Z digest=sha256:8318f8e9645455a25d28bf6cf0d857e42f4a1a746b71b45b96fc66c352b5b0e1

Observation 391030fa-4064-41f5-a09d-32a19b9a97b0 · outbound

This paper cites Conflict-averse gradient descent for multi-task learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Conflict-averse gradient descent for multi-task learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.001905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.436638Z digest=sha256:92bd9132bc0c91db3076b89a3bb3d628b9fe73a9b8865b9b5e8ec8678d149cb1

Observation ff021dd4-6f61-420c-ad80-4985741505ef · outbound

This paper cites Moka: Open-vocabulary robotic manipulation through mark-based visual prompting.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Moka: Open-vocabulary robotic manipulation through mark-based visual prompting

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:04.743852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.496763Z digest=sha256:bd33502d3a18544d88b3f42e56673849544c010ca1bffc2bf2430b119a6139e3

Observation 91825c01-af53-4a35-827c-03fa771c2fe6 · outbound

This paper cites Scaling laws for fine-grained mixture of experts.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling laws for fine-grained mixture of experts

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:04.481489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.558081Z digest=sha256:456ee0e2342704bf1c76f62d9b2509ff856baf541dea15c8ce3e22b6d68fde9f

Observation a5446eea-762d-424c-bc91-e942572d7a78 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.612205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.612205Z digest=sha256:2d448f40f1cd34e7322d4a81887da1796a5087bcd30b7603d8bc69b9669acb0c

Observation 2a72955a-105a-4fd6-824e-1a6f81559308 · outbound

This paper cites An Empirical Model of Large-Batch Training.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners An Empirical Model of Large-Batch Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.670391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.670391Z digest=sha256:04a6f8e27a0295324d448e6eef91789480aa53e6d655ea3a95ca9877d3df52c4

Observation 0add11d2-7c60-4e76-9f95-c193c5bd944f · outbound

This paper cites A., Veness, J., Bellemare, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A., Veness, J., Bellemare, M

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:59:04.257681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.729363Z digest=sha256:359e101bd392532b07b7d4b3430515afbd6aa3b1b9fdf9d5a15f252c0f7b61fc

Observation b5a37fcc-0fb6-4ddd-b758-ab4199734896 · outbound

This paper cites P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.997563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.786249Z digest=sha256:1602c945b5feda02e01cc5bc257a63a97d2b6ea347ad97c032ac44d2243c4198

Observation abafc6a0-52de-4d8d-a1fa-c3a640fc15c7 · outbound

This paper cites Tactical optimism and pessimism for deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Tactical optimism and pessimism for deep reinforcement learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.761137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.835032Z digest=sha256:877819ef100d7875b6ed9ef01ff4eb9ea19861a67058815b7f8e1b47f33d71e8

Observation 8f320d94-7ca4-4ed6-b067-77c05124fbfe · outbound

This paper cites and Cygan, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Cygan, M

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.545410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.913669Z digest=sha256:27bc9706d254bdd003516457e3a736a6477ba97e14495b445133dcd0b817022b

Observation 19548d37-4bec-43bc-af16-66f226479d82 · outbound

This paper cites Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.075651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.075651Z digest=sha256:bdf8b74b66b29e640284c5389fb6890fde1d38a93c33eff9837e16249df8a25d

Observation b79eba20-6d7e-4ddb-b599-54112b534780 · outbound

This paper cites Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.287792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.287792Z digest=sha256:683f8f8fa40daa26311b71a4e113d8de6b94520bcec7c8132016b9eeb7e7c52c

Observation cc34fa6c-1f32-40c8-b836-35a24ae65bc6 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Dinov2: Learning robust visual features without supervision

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.335629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:51.408062Z digest=sha256:b3a29c0a184d8316b34a9ca094b1ee0a92a99ed8698ac63e8fad1a5b86d75ebf

Observation 08140eb0-776a-4c7a-be59-1307dba4a575 · outbound

This paper cites Training language models to follow instructions with human feedback.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Training language models to follow instructions with human feedback

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.565324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.565324Z digest=sha256:85161c1e7bba85c6d7a40e4f9cdd152e1d889f2d00d864d659ed261ff2101b2c

Observation 3bba43ec-1eaa-4744-9ece-92d7769e9d23 · outbound

This paper cites Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.719632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.719632Z digest=sha256:dc758cfd56abe10949352b2859016bebdbf8e149c03e6508f4d2a2a0ebd335e8

Observation d4eaee66-7207-4e20-a551-212ec8395e2c · outbound

This paper cites Is Value Learning Really the Main Bottleneck in Offline RL?.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Is Value Learning Really the Main Bottleneck in Offline RL?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.847608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.847608Z digest=sha256:7e544ee7ee56deaf1ec8c37a67b344f8987ab4c4335d84286b46f7122a10a07a

Observation f4dbe154-c0a7-4194-aec4-404adbed9de7 · outbound

This paper cites Flow Q-Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Flow Q-Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.955281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.955281Z digest=sha256:2d451111e09a4daa1e0477d6acfa26144b0ceab99f37367be57f769ca8c07b1d

Observation 8b01f985-6863-4184-b4fd-d48391dbff62 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.070645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.070645Z digest=sha256:c814b86f275a9af40b8a6129d73d11a45687eafc2726fc73cebbb27c169b27bc

Observation 5c9ea69e-97fa-4277-973e-b05b52d26d1f · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.113363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.157484Z digest=sha256:43489e47aaf8ee5156d0824fe23a499167704eaec30a68fb0c2dcf050d96c4d3

Observation 1f3c6d5b-3fbd-431d-b5ed-0507340cfac5 · outbound

This paper cites Efficient off-policy meta- reinforcement learning via probabilistic context variables.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient off-policy meta- reinforcement learning via probabilistic context variables

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.900845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.275833Z digest=sha256:1f63118a48a5b8213fbf05c6f87fe5ef6de08297770fb003168ea0f426223737

Observation 0935f476-d90e-4a90-ab89-f9845ff239e6 · outbound

This paper cites A Generalist Agent.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A Generalist Agent

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.357891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.357891Z digest=sha256:dfba98b8d31491cb819c8df785ab2f957fd0760f3865b1d7cef0d938ac87f8ec

Observation fe4849cb-8398-425d-b46a-d6a609b6f0a4 · outbound

This paper cites Policy Distillation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Distillation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.473229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.473229Z digest=sha256:086768011e59354f04e02db7eeafedd4bcf2e0233534b72c990b04f2601d8544

Observation c54498a5-69f8-4676-9eeb-e9dc48e64daa · outbound

This paper cites Value-Based Deep RL Scales Predictably.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Value-Based Deep RL Scales Predictably

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.599976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.599976Z digest=sha256:9772532c66ae4dcd779b4990d81dd4aa277305bfca95d932e3d1eef633b78dcb

Observation 312c1e26-ea4e-4218-b376-37e15e0b2dfc · outbound

This paper cites D., Courville, A., and Bachman, P.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners D., Courville, A., and Bachman, P

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.700604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.687060Z digest=sha256:61ab43b07d69710a9f3f61faa6508166ba2544c459dff4c7d83cb6f4134f46dc

Observation 2bf15b1f-6df2-48bb-9ed5-19c29a58c920 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:02.436833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.799338Z digest=sha256:9cb8da9535ce3dd2c0b155400c210c9244a6f62ddd2174f2b142a74435a0400b

Observation 980db95b-203d-45db-9dc7-f0f6d759858c · outbound

This paper cites and Koltun, V.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Koltun, V

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.224585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.948108Z digest=sha256:fb7081bf46130e47a871b5cedea705450cd0a9035074178e9bba7e071abade84

Observation 26518396-25da-4193-9b3b-c8c1a5ad8717 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:53.082154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:53.082154Z digest=sha256:56e311d88f77b81f78c59e5dcf193d046a05b9d36b0e9b5a8edd5b0f915aa958

Observation 2d99e650-1b44-4599-ae47-53f38f3ab586 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:01.908957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.239036Z digest=sha256:d227b8fc364764ea7865e7a911513398c1252a96a2cc509cd15e0b0ad1a7ebd0

Observation d63ad454-a067-4c4b-8a9c-4c038b1d35fd · outbound

This paper cites Deterministic policy gradient algorithms.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Deterministic policy gradient algorithms

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:01.644985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.386932Z digest=sha256:e5b58df28420bf0a7e83d297636721e2e038ffd506419c165a82759b997cde3e

Observation 6390f0ca-ea85-4fef-9b4a-b50d88db26b9 · outbound

This paper cites Mastering the game of go without human knowledge.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Mastering the game of go without human knowledge

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:53.487874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:53.487874Z digest=sha256:de5b3a852dc6ba8c466ee8005e55cf46b0c61d21c930727b5ed5620a182867a5

Observation efe21a7b-af83-4aea-b805-06006d46249a · outbound

This paper cites Multi-task reinforcement learning with context-based representations.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Multi-task reinforcement learning with context-based representations

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:01.310812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.631167Z digest=sha256:5fb1b5eb676b04ecd0ee9de9877c1dc3587bff23f234e910c45d8b4b88c49f02

Observation c9aa705f-9b4b-44d3-a5c5-4b0ef5788f1a · outbound

This paper cites T., Abdolmaleki, A., Zhang, J., Groth, O., Bloesch, M., Lampe, T., Brakel, P., Bechtle, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners T., Abdolmaleki, A., Zhang, J., Groth, O., Bloesch, M., Lampe, T., Brakel, P., Bechtle, S

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.982770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.780565Z digest=sha256:c5472c971d565260239ef96397fe4e29c80c977b98886d49b7eeac022816d368

Observation 517bebf6-d433-4734-b018-4b8d1cc626b7 · outbound

This paper cites Paco: Parameter-compositional multi-task reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Paco: Parameter-compositional multi-task reinforcement learning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.682511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.893424Z digest=sha256:fcf6c6e61321214245a7cbe89234fe4f8d52016c26e1d4b99ea5d044255391c6

Observation 3638475a-9732-4c67-a418-49cf4febf075 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.004334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.004334Z digest=sha256:c106e066d5cfd0e320654900c4225e1c05a00987c1dc1acbf2e526691a5a01b1

Observation f4797f4a-225b-42be-bab4-c83fa7fde7b1 · outbound

This paper cites DeepMind Control Suite.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DeepMind Control Suite

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.169006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.169006Z digest=sha256:43b76ca45d3a8a353fd2c0fdec262fd2c51aaeff9578eadfe38a6e2bf11a2b03

Observation 9ad56e27-cc51-49c5-973b-9397858f3258 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Gemini: A Family of Highly Capable Multimodal Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.270164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.270164Z digest=sha256:8b33d4dc9f5ae97eb7a8c75589efc655d60df1c477bed3883d56265c8be3316d

Observation 13c97a2e-7c15-4072-ada9-4c0f675e3410 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Octo: An Open-Source Generalist Robot Policy

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.408075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.408075Z digest=sha256:9801d34deabdd465dc5185a0e9bb0c0dba2df3ad648beb1f9f683d0dad4eeac1

Observation 4abc9372-2b55-4050-b912-923702c19438 · outbound

This paper cites M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.375399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:54.542081Z digest=sha256:6e5e81ca490a391aab2e6bb844f4cbca84518a95b07c5d0df6feac6b3ccb5046

Pith citing papers

Observation bc77ead2-1f30-42fd-9916-5280d93a49ee · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:05.516882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:05.516882Z digest=sha256:816ebefbaa523b11e3d35a5ad6b7019d81d106b2cf33164564606a9431e1ae5d

Observation 0b865dce-dccd-424b-8539-56b73b144627 · inbound

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning cites this paper.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 1937

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:50.060301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:50.060301Z digest=sha256:f559501eeaa4d3b465339931da38039209a5487ff97a9cffe3fb2a2c8fc2d97c

Observation 60621c76-8d7e-4e50-99fc-2dba385bb7b0 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.973880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:bfdf65d07f717651193f8ba12584284201594ba6b0710c2ced1afe34042b7b93

Observation 888cc10a-ff82-4837-a3e4-b20e78351f66 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:49.716390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:685b24d1977c7fce80ceaef445da6ed566b3bf169572390ae73d7c6016fd3369

Observation 9c5a25c9-0ec3-42b4-81d1-2f0be6bf0658 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:12:41.363803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:2f81fecdf3559f52ad7cae44b5eb8665ca14999fcce6065dc0ab3e8d71131b36

Observation 309e830e-b3ea-45f2-9ca8-588e9c61e2ee · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.607354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:33:19.038889Z digest=sha256:9fea8f9fed4e3c70733a6c6fba2a3faa38621eacbf87e6334ba01457d6285e95

Observation 2664d873-3fe7-4957-8bac-b87353feff58 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.804374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:27:38.643667Z digest=sha256:e629c799250e8e832be384ed8a0b86029b5a50d7ff84925e736f4f02cf9bbe8b

Observation 85285cc0-7eaf-4f1f-be3b-64a47fdd07a0 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:09:12.743772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T23:04:12.943222Z digest=sha256:b49f2ca5bfbbb6e8d9f70d67279770a3289d1a204584757a0d3514bf2e070aff

Observation 08c63a93-2769-4044-b758-df2313d4dca4 · inbound

Debiased Model-based Representations for Sample-efficient Continuous Control cites this paper.

Debiased Model-based Representations for Sample-efficient Continuous Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:12:28.425885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:09:01.132424Z digest=sha256:9d6ce4d1ca18c5492fee160b9d3a104a99224dd1d9b754de49226cfb956d8970