Pith. sign in

Paper Citation Record · LEDGER

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

As of 15 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 10 inbound Pith citation observations for arXiv:2505.23150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23150 v1

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:54.542081Z

measured 110 of 110 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:48:47.358278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T23:09:12.737900Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved76
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c403830a-7572-46f5-8a4c-4e8f99131eac · outbound

This paper cites GPT-4 Technical Report.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:44.854747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:44.854747Z digest=sha256:3a375f17a48e788ccb89c53c56af63f89e9ce1c20ee5d6b0c144db9288d49848

Observation 80c9da86-63a7-4a64-b95f-3a0d48e2c8c8 · outbound

This paper cites Provable benefits of representational transfer in reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Provable benefits of representational transfer in reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:44.967526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:44.967526Z digest=sha256:4179fff1b8df7686e291990278a8857413aaaa889d4e63cae672c7c891a70132

Observation 713717d6-e6bd-4c9d-96bd-c26975730a9d · outbound

This paper cites S., Courville, A., and Bellemare, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Courville, A., and Bellemare, M

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.019350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.019350Z digest=sha256:221b58c583188b66fd8db8b83177a96c7548f2a70d2b6ee2587a399ed967d76d

Observation 97a075bd-184a-4409-aee0-3305870ecba4 · outbound

This paper cites S., Courville, A.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Courville, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.121469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.121469Z digest=sha256:23c2ec2be1c9ba74faca76f4cb4af996e2524666fd5038ec2643bb70aa2c031a

Observation 7da3fa4c-e85c-4342-99ec-4195f8b9d31f · outbound

This paper cites Hindsight experience replay.Advances in neural information processing systems, 30, 2017.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Hindsight experience replay.Advances in neural information processing systems, 30, 2017

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.197775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.197775Z digest=sha256:5023511be100d5aeb764ee44100e6ee3d2302df119debfbf2ddf8b1927ebdea6

Observation bf984bb8-9e57-4b9d-b244-bc6220398bcc · outbound

This paper cites M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.269538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.269538Z digest=sha256:f65a342ff507f24f288cbbf4fc612afd57593a976e9b05a13db47b32e38ba785

Observation 74c0d758-659e-43e9-ab11-00f55115d6d5 · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.345247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.345247Z digest=sha256:dba1bdece42e8be9f7cb3a2f2765bc8ca15c465f1c0282488fb91d86b0f52ec2

Observation 98169282-9aa4-4dae-8152-7ca68d2c59fc · outbound

This paper cites J., Smith, L., Kostrikov, I., and Levine, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Smith, L., Kostrikov, I., and Levine, S

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.433639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.433639Z digest=sha256:5d209c42fd7017eb86d9082fa859ad233cfea2fef799cabbde8c45a9a3643b0a

Observation 3b1b6998-d6db-4859-ba31-625e5dc47236 · outbound

This paper cites J., Schaul, T., van Hasselt, H.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Schaul, T., van Hasselt, H

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.571178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.571178Z digest=sha256:0c068d9c3211a8d650b63cc2f9a782635d9513d898ad1e740f30f955daa771e7

Observation d9949fcb-109c-4545-853c-02164a0e7727 · outbound

This paper cites G., Naddaf, Y ., Veness, J., and Bowling, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners G., Naddaf, Y ., Veness, J., and Bowling, M

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.663223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.663223Z digest=sha256:b040e1744ee96ab63387b500828a89297bc95d64964f9ce1a72d14cd12ddab67

Observation c5f64815-4508-4e5a-ac03-e3cbf631d284 · outbound

This paper cites G., Dabney, W., and Munos, R.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners G., Dabney, W., and Munos, R

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.750044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.750044Z digest=sha256:dd15be6420317d47390ed7be9ee6337e57ea5135c2c97697b6431a4aa8f17c37

Observation 08d1053d-46cb-4f71-880d-78bc3dd62be7 · outbound

This paper cites Dynamic Programming.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Dynamic Programming

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.804154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.804154Z digest=sha256:c7b88655a9f726ad8191e874dc9c8ce96443005c5087fa03ca291a3cfb266d69

Observation 2486c5fb-c6b4-443c-8cdf-fd42dd633eb4 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.874874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.874874Z digest=sha256:58f21e0e7a81e48a854766f3a3db2602f7e01ebad4121525a4776e017ff2aac6

Observation 93b4fe42-9f73-4d76-82fa-1430bf502f95 · outbound

This paper cites P., and Weinberger, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners P., and Weinberger, K

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.968173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.968173Z digest=sha256:967d4d4d0845985d7ed1ae11191796f763f2b27a8d79024888b86c46d1a4e0dc

Observation 68d41d36-8aae-423b-996d-11f1b1fdfef7 · outbound

This paper cites J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.065866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.065866Z digest=sha256:353d57132d64287c8f70dc4f5c4803dcd3458d5c48b77068c325f6e2e4aaf16b

Observation 63c3a67c-9aca-4c18-9056-0c246c26dc54 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners RT-1: Robotics Transformer for Real-World Control at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.158225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.158225Z digest=sha256:7a5c62fd6fe8ac60e8b236ffb16141209c2255124c01f5d40c43b758ab734641

Observation 778f8cbc-d862-4bf3-b45f-d40d66689849 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.274336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.274336Z digest=sha256:427fddf8605e70cf6717246187afc5143620d56f5af9e9049803e26705b9f578

Observation c8a57e69-4a22-4384-9865-e7bbadf3fcad · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Decision transformer: Reinforcement learning via sequence modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.379002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.379002Z digest=sha256:f19eabd6c810213d1b7451d20d6da8e22d6a40bfdce8afe1827e1324c3321f38

Observation 8bfdf1d4-ebbf-4fa2-969a-3b77060d52cf · outbound

This paper cites A system for general in-hand object re-orientation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A system for general in-hand object re-orientation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.445275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.445275Z digest=sha256:b2862274a543c8871df8aee90a64614e45d457d495d1c03700c67ea0c462aa6f

Observation 5c71d3a6-7cbf-4467-85f6-f06c0eb2f829 · outbound

This paper cites Gradnorm: Gradient normaliza- tion for adaptive loss balancing in deep multitask networks.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Gradnorm: Gradient normaliza- tion for adaptive loss balancing in deep multitask networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.566082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.566082Z digest=sha256:27a293debd5fa0a9ba504131bc26b40e8b71fe391e68e121b944ee01048d8f37

Observation b133e83c-fc66-447e-8f32-286d8c5f2621 · outbound

This paper cites Just pick a sign: Optimizing deep multitask models with gradient sign dropout.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Just pick a sign: Optimizing deep multitask models with gradient sign dropout

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.662012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.662012Z digest=sha256:18cf51be865563eabd8cab58f5eaf8037c4f3c1dd851224d90e7bd382e177350

Observation 725d0859-c737-4840-9e5f-ba3f83d84306 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.707628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.707628Z digest=sha256:927ee33659b00435cfbdac0269389194ed6cbfaee112ae5fd09bfc590cfcf02c

Observation e5752454-cbfc-4efa-9bae-76e8fa6c8afe · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.750339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.750339Z digest=sha256:a102a30dfca694a275756182202ab3f0c404cf26f4957bd978193a07946349cd

Observation f173b2ca-f1b5-45e1-a98b-0ac0ff439627 · outbound

This paper cites Better exploration with optimistic actor critic.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Better exploration with optimistic actor critic

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.835161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.835161Z digest=sha256:162555cb94c06688a92cdaa5243b1bc9ae1d00fd70696d7c9a4d56ee0c862008

Observation a9579dbe-bdc1-4982-8b39-ed18dca7ba0f · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.867259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.867259Z digest=sha256:a35eb6181b1faa43efc62d0806129a53b328c657276a47f8f12e37823ee580c2

Observation 544c6b33-a6c6-45fd-ac93-60cd1d1f1f1d · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.915785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.915785Z digest=sha256:3722ca09bdec9d399f66f2b0ff9c335c2a59bd81aea40113cd3de88341477274

Observation cedc50c4-c18d-4625-9599-87af9b1a4611 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners An image is worth 16x16 words: Transformers for image recognition at scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.009204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.009204Z digest=sha256:88db0f81c83989541d4b55b9a4b676617e34a93d8b4acda779a4aeb89a7bbf75

Observation e88b3017-0087-4ffc-864a-353bd7bbd744 · outbound

This paper cites S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.060172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.060172Z digest=sha256:1b2697b1f6bb52d493d88d9a4d5412cbeb65aa7e306eb1353a9b6a1e9369afb4

Observation 8156fac5-839e-4306-9452-eb846bba34f8 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.128470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.128470Z digest=sha256:4f249d6cf0ae5bb14ba857382e4f081c7c652da93573575e36ccb2599960438b

Observation a990062a-6dba-4f00-b38e-c2f4518b63d4 · outbound

This paper cites A., Chebotar, Y ., Xiao, T., Irpan, A., Levine, S., Castro, P.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A., Chebotar, Y ., Xiao, T., Irpan, A., Levine, S., Castro, P

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.227015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.227015Z digest=sha256:403778fdc65a1412c5c638204909b35a98aaccbda7c2ad31de665bc7707ebf7e

Observation ba3cede9-4995-41b8-a536-f8b58699f1f3 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Model-agnostic meta-learning for fast adaptation of deep networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.334044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.334044Z digest=sha256:a19be79842d2a55d5785a480cefc45e4707bb08ba1612ffb87d8925488046d49

Observation 5e78924f-1e27-4136-9ecb-d46590b3ef89 · outbound

This paper cites and Gu, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Gu, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.398149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.398149Z digest=sha256:92c48dc0664719a28eb8eabccc24122206bc4d24b4820b499be3667b36db835c

Observation b3f8af50-5032-4c35-8ef3-b0e167b5c22f · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Addressing function approximation error in actor-critic methods

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.450522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.450522Z digest=sha256:db70c8f99166ae1fa9de3ab207af7701a956f7e5776d3ec6438bda9bd4a2a46b

Observation ad6cab91-fa5f-41a7-bd2b-995743cf13f6 · outbound

This paper cites S., Precup, D., and Meger, D.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Precup, D., and Meger, D

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.514281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.514281Z digest=sha256:91f7cc7e8ba53b15e29074eb5eb046b3234115265e03c50a270df962af3e373c

Observation 3f5c7ad6-7ec2-46ff-b60d-2c762b03bdd6 · outbound

This paper cites Divide-and-conquer reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Divide-and-conquer reinforcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.599219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.599219Z digest=sha256:12566e5be6f502a78acea80ee09bf6fec14222f3c70d2a7320becdffc0121461

Observation 11695065-eb2b-43bf-8acd-b1b64740f6bb · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.702018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.702018Z digest=sha256:07155f7760004ed9a0a9d82e691fef337daee75aa155763e5c0d7f299603f9eb

Observation 9b11e6db-448c-4902-ad8b-699cd749dc7e · outbound

This paper cites Mastering Diverse Domains through World Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Mastering Diverse Domains through World Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.810738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.810738Z digest=sha256:1ef5ed7ff1235be67a68886da606f1bec7cb1e63a87e16c8ec2ce4a93db5857b

Observation dcd824f7-6396-4ee8-87f5-7728212e1b21 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.926005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.926005Z digest=sha256:bbc522c26d60490404361d752bdf09160398a699033cf974a8460c0c018a5fba

Observation ee899869-f4ac-402a-bf1a-243769f12730 · outbound

This paper cites R., Millman, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners R., Millman, K

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.989117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.989117Z digest=sha256:dbd063f675fe08935a8ab76aac71686dd0833efafcaf332cfca33fc29fc703aa

Observation 83679e9d-c067-4d62-96aa-c79d0a711e38 · outbound

This paper cites T., Wang, Z., Heess, N., and Riedmiller, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners T., Wang, Z., Heess, N., and Riedmiller, M

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.058887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.058887Z digest=sha256:d4665c46f8114090eaeb250e3dfd32a477e802bf7dbc656abd9f766d9c63f44e

Observation 0e9bb63e-08e0-494c-bd15-61e1e3d3a57a · outbound

This paper cites Efficient multi-task reinforcement learning with cross-task policy guidance.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient multi-task reinforcement learning with cross-task policy guidance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.115400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.115400Z digest=sha256:8427f4c58c09af0987df3722eba5170d53a56bad7d17eae516600e4db59a4a8e

Observation b8e672c4-b654-4d62-9126-c99e3324d0bc · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling Laws for Autoregressive Generative Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.198228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.198228Z digest=sha256:4191385d8b4f6ae50a7b8b2b1faf373875daac87527f9cbff944d471e139dcb8

Observation aa26013b-9414-44dd-8134-fb1b1af1cdd2 · outbound

This paper cites Multi-task deep reinforcement learning with popart.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Multi-task deep reinforcement learning with popart

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.306053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.306053Z digest=sha256:22bbe3afed919e060897a64db433f1c220c4987a3726b81fa78b254676312489

Observation e10abf86-0652-4874-9896-dc073ebdf597 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Training Compute-Optimal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.423350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.423350Z digest=sha256:94ed679775ef9d3344f02e9cf352cd444a98a1c84ef7e5a7ff9bc05671d8c0c5

Observation da2ff812-3cf0-4a19-a9f0-32166cf81636 · outbound

This paper cites Otter: A vision-language-action model with text-aware visual feature extraction.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Otter: A vision-language-action model with text-aware visual feature extraction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.535699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.535699Z digest=sha256:d0433dd654617d0e314d9088ae9debd1b88b45302673a6835c49267a1e5d97e3

Observation 1d70b86f-53ea-46b7-86c0-9f050d402e4c · outbound

This paper cites Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.664444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.664444Z digest=sha256:a653ab8367acf2e3e65bc6f7e4525fda42f736297872eada101831d4debb7d17

Observation 433abc24-6485-4cf0-9535-847bfb6ac080 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.794678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.794678Z digest=sha256:4d89f015698e80a63e2abe4f0ab1dc830c72c0110ecffa0d9de8cdc40c3453f0

Observation 1c78482a-fe9b-43f3-9780-a093bedcb51a · outbound

This paper cites H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.510973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:48.918826Z digest=sha256:1b6bec58dc3ee9f40143bcf8f356c0256539538c95b27e6e1222b48a71ebfac4

Observation a618c6a6-b9fc-42b8-81af-bfefb5440eaa · outbound

This paper cites Scaling Laws for Neural Language Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling Laws for Neural Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.042717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.042717Z digest=sha256:e50700a56be090fe8f8b65bfda26d535b23a474ab64780a09755355a35dba0b0

Observation 5a2632db-b1fc-431f-b978-3e0be2c69394 · outbound

This paper cites Champion-level drone racing using deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Champion-level drone racing using deep reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.350502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:49.165271Z digest=sha256:58a421679764073aca80b3becdaf81f59e5f1df3a6e8e4dbe8ed36ce010601cf

Observation 9dd9a88f-9958-4e08-a061-06a00151d3e1 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners OpenVLA: An Open-Source Vision-Language-Action Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.306614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.306614Z digest=sha256:9de9005166d84628853b57108a260264fc5d88cbabb4176ce5b2e72db9f46cfc

Observation 82ba696b-a068-4fdc-942f-828e6bd8b89a · outbound

This paper cites and Ba, J.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Ba, J

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.176243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:49.440203Z digest=sha256:b6b12d8168a5a87ce0e5abee9f5a81d3622c6c4251c96746b01343cf39bcdb85

Observation 25391644-5245-47ae-a423-bd5dfe19e699 · outbound

This paper cites C., Lo, W.-Y ., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners C., Lo, W.-Y ., et al

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.005942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:49.554910Z digest=sha256:00d4804ba761c59440c94ad416dc41e9ecc1cee178cb70ecffe8b69afdf0b99e

Observation dabe2b8f-9758-4483-88a5-a239cd795e92 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:05.816415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:49.704891Z digest=sha256:cae686cdedab67dd36affc288b719e580ed48eede9e421d3624ac56b165e802f

Observation 0b377d3d-6be3-4aa6-b19d-e7f2245a7783 · outbound

This paper cites Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.816226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.816226Z digest=sha256:ac16788b81ae0a13ed472c5cc8c58fe07a2aea5307b2a8863b2d366fcda7654f

Observation e0fce54b-56f3-4e46-9d78-a750c2e7257d · outbound

This paper cites DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.941083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.941083Z digest=sha256:83aca1aa82df3cf2ad72c552ffee704c565c56f59f430dd33703f13211388bd4

Observation b92fdb5d-b34e-42b9-9dd8-0b79474c4011 · outbound

This paper cites RMA: Rapid Motor Adaptation for Legged Robots.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners RMA: Rapid Motor Adaptation for Legged Robots

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.021389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.021389Z digest=sha256:85c5cce6dde7f1160c48c528940d80e786a6a032387695d11bf4bbcac186dd99

Observation b7b884ea-e9b5-4d08-8c9d-2d6a9dafe2a9 · outbound

This paper cites Offline q-learning on diverse multi-task data both scales and generalizes.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Offline q-learning on diverse multi-task data both scales and generalizes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.630424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.084622Z digest=sha256:84c9f673e2780057ea8196e0e44561a78fbd36cf4125439a6d534e6cd88c8bb1

Observation e53f8be0-5c13-4972-8dd5-fe5ce4b48e33 · outbound

This paper cites Reinforcement learning with augmented data.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Reinforcement learning with augmented data

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.433016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.148267Z digest=sha256:38de0b6cf42e4035fbea5662890b57aa2200011f9e12c6539528d72990fa6ede

Observation 6e598f6e-02ff-41e2-80f0-7de53dca1f64 · outbound

This paper cites SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.224254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.224254Z digest=sha256:ab2d990e3a030bd1bfdfe1e8f566b69000e4b8778b94e3f0a5335f82d0737e47

Observation 0ac955de-f366-46dd-950e-f7f7e1d1f34f · outbound

This paper cites Hyperspherical Normalization for Scalable Deep Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Hyperspherical Normalization for Scalable Deep Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.269251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.269251Z digest=sha256:3664bb74468376bf162fa4887aeb43ad679bab26d1390ccb5bd746a5bb21917a

Observation ae14b0df-74e1-4d93-bc6c-6e594c8ee1ac · outbound

This paper cites End-to-end training of deep visuomotor policies.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners End-to-end training of deep visuomotor policies

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.163798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.321875Z digest=sha256:48ff13c17b0dc19c58d40b70584650655897518d7c6a8086f8ea8298a4109d6c

Observation ba682875-3697-4628-a572-3f871e74de31 · outbound

This paper cites DeepSeek-V3 Technical Report.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DeepSeek-V3 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.379474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.379474Z digest=sha256:0ec48a701e2b676046261a80bd99356cd9e6daec53463cf234ef08be1295866e

Observation 391030fa-4064-41f5-a09d-32a19b9a97b0 · outbound

This paper cites Conflict-averse gradient descent for multi-task learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Conflict-averse gradient descent for multi-task learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.001905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.436638Z digest=sha256:76b8c474d55e3a5c3510e388b8cb5b818cfcadbd5b520ecbe4f514e896bac4cf

Observation ff021dd4-6f61-420c-ad80-4985741505ef · outbound

This paper cites Moka: Open-vocabulary robotic manipulation through mark-based visual prompting.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Moka: Open-vocabulary robotic manipulation through mark-based visual prompting

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:04.743852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.496763Z digest=sha256:8dab066fe6055de67f6e06ae70d88b61b018cb81d662d55a346c0b136dccdf56

Observation 91825c01-af53-4a35-827c-03fa771c2fe6 · outbound

This paper cites Scaling laws for fine-grained mixture of experts.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling laws for fine-grained mixture of experts

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:04.481489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.558081Z digest=sha256:279eaade1e2e13db8a717d73528bdf1196cb093984c11190b8eca331d45dcd82

Observation a5446eea-762d-424c-bc91-e942572d7a78 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.612205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.612205Z digest=sha256:e788c91c60161ef0c24aa3667b9d158ce2f0fcf598c14d8ac3958812e170b26f

Observation 2a72955a-105a-4fd6-824e-1a6f81559308 · outbound

This paper cites An Empirical Model of Large-Batch Training.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners An Empirical Model of Large-Batch Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.670391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.670391Z digest=sha256:c594a958d59d6c9d2bfaadbddeaf485379d232768ad5345499dc60d3be764a7b

Observation 0add11d2-7c60-4e76-9f95-c193c5bd944f · outbound

This paper cites A., Veness, J., Bellemare, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A., Veness, J., Bellemare, M

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:59:04.257681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.729363Z digest=sha256:88eb738d14b619a2bc00a7aef2ccbf0f381aa31d45e085e643a1c76194c541ce

Observation b5a37fcc-0fb6-4ddd-b758-ab4199734896 · outbound

This paper cites P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.997563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.786249Z digest=sha256:bec14212cfbf7b3dcfb728c505774cc307663cf2bb6d404e6bfde8cc9f388844

Observation abafc6a0-52de-4d8d-a1fa-c3a640fc15c7 · outbound

This paper cites Tactical optimism and pessimism for deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Tactical optimism and pessimism for deep reinforcement learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.761137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.835032Z digest=sha256:11892ba9744b1a37834fe31b2a6a52658a07087b3902e5cecec7ab860913bb65

Observation 8f320d94-7ca4-4ed6-b067-77c05124fbfe · outbound

This paper cites and Cygan, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Cygan, M

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.545410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:50.913669Z digest=sha256:c6a1507b25d1305cf27ec8161e51c5ced123d68645cf47be37ab04107a84ac08

Observation 19548d37-4bec-43bc-af16-66f226479d82 · outbound

This paper cites Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.075651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.075651Z digest=sha256:9f3ec64a5a71a38312ed44c15d8459044d43e93577735c69e1d44833d9cd5858

Observation b79eba20-6d7e-4ddb-b599-54112b534780 · outbound

This paper cites Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.287792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.287792Z digest=sha256:5d72f410a71e167ea7c0321ca515dd73f115af0f4038da5b7bbfe32de4c197aa

Observation cc34fa6c-1f32-40c8-b836-35a24ae65bc6 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Dinov2: Learning robust visual features without supervision

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.335629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:51.408062Z digest=sha256:7d5991e8860e6ce5d60230394f9e00735b8e60776573ec1b3c91418f2190378f

Observation 08140eb0-776a-4c7a-be59-1307dba4a575 · outbound

This paper cites Training language models to follow instructions with human feedback.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Training language models to follow instructions with human feedback

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.565324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.565324Z digest=sha256:796c1c1f9c7dc4bd75e8f6a736f94de9af8dd73cbeb371d9636dc48a09995c0b

Observation 3bba43ec-1eaa-4744-9ece-92d7769e9d23 · outbound

This paper cites Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.719632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.719632Z digest=sha256:3b6637e9a6258cb1685474a94b973fe981ad94fc0e4b356bf0c4436b7de41b1f

Observation d4eaee66-7207-4e20-a551-212ec8395e2c · outbound

This paper cites Is Value Learning Really the Main Bottleneck in Offline RL?.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Is Value Learning Really the Main Bottleneck in Offline RL?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.847608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.847608Z digest=sha256:6cc943d477fce4d5849b6b1daf01851fb9f30190e2a0b66359318db37b8cc16d

Observation f4dbe154-c0a7-4194-aec4-404adbed9de7 · outbound

This paper cites Flow Q-Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Flow Q-Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.955281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.955281Z digest=sha256:617215e39a8f7a2dc4b3a18e8d73aca4b809914b842c2a33641a03d47f4d4ef2

Observation 8b01f985-6863-4184-b4fd-d48391dbff62 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.070645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.070645Z digest=sha256:e8b750f87d5ff0d4c559c90ff215b500bc00563bef679d6be6522079591b8fa9

Observation 5c9ea69e-97fa-4277-973e-b05b52d26d1f · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.113363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:52.157484Z digest=sha256:0cb943eecd3fafe54fb2fe08de6b584ca3de3984aa7aaac41dd0f065913bc869

Observation 1f3c6d5b-3fbd-431d-b5ed-0507340cfac5 · outbound

This paper cites Efficient off-policy meta- reinforcement learning via probabilistic context variables.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient off-policy meta- reinforcement learning via probabilistic context variables

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.900845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:52.275833Z digest=sha256:76158a2d07188c626908f57d5e9a8f4b992434cb18d0613cc2df26026748d45c

Observation 0935f476-d90e-4a90-ab89-f9845ff239e6 · outbound

This paper cites A Generalist Agent.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A Generalist Agent

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.357891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.357891Z digest=sha256:26bedbff320aa7546344df988e07538fedaa6af813f3b377ffd2088d94038149

Observation fe4849cb-8398-425d-b46a-d6a609b6f0a4 · outbound

This paper cites Policy Distillation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Distillation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.473229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.473229Z digest=sha256:cfd95a4c3056c81a2fd4969381ea903accb0424dae0cefc7de5553b728d12b16

Observation c54498a5-69f8-4676-9eeb-e9dc48e64daa · outbound

This paper cites Value-Based Deep RL Scales Predictably.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Value-Based Deep RL Scales Predictably

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.599976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.599976Z digest=sha256:3d8ad848ea790430af371ac53217261cc1c7e4fe5b4f2b5894101ee74e1ca06b

Observation 312c1e26-ea4e-4218-b376-37e15e0b2dfc · outbound

This paper cites D., Courville, A., and Bachman, P.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners D., Courville, A., and Bachman, P

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.700604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:52.687060Z digest=sha256:082eb63e65c697fbb277c9dab473bb3c89f3f8dac083ee98609873a4ec317f33

Observation 2bf15b1f-6df2-48bb-9ed5-19c29a58c920 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:02.436833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:52.799338Z digest=sha256:27787dfa6d9be9d7f58f644880edf97051f91a23ea3f42bab7de78ab1ad7ba21

Observation 980db95b-203d-45db-9dc7-f0f6d759858c · outbound

This paper cites and Koltun, V.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Koltun, V

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.224585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:52.948108Z digest=sha256:d2457dc34e39e2f241c5c895de2dcc70b4fedf6f521c44ab06f3fdab6197d91f

Observation 26518396-25da-4193-9b3b-c8c1a5ad8717 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:53.082154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:53.082154Z digest=sha256:7a7d8da4ddf745a68c3584ebfe0637ce2f3881507c9177138b03d695ae796de5

Observation 2d99e650-1b44-4599-ae47-53f38f3ab586 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:01.908957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:53.239036Z digest=sha256:6d789d5557d698e50f339eb6874a62089b527ab417a798fe1f45a81e8b64b801

Observation d63ad454-a067-4c4b-8a9c-4c038b1d35fd · outbound

This paper cites Deterministic policy gradient algorithms.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Deterministic policy gradient algorithms

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:01.644985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:53.386932Z digest=sha256:cfab1f043130e460ec5915a939dd239eee38af8f26bafc5cadb1b9f220349fcb

Observation 6390f0ca-ea85-4fef-9b4a-b50d88db26b9 · outbound

This paper cites Mastering the game of go without human knowledge.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Mastering the game of go without human knowledge

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:53.487874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:53.487874Z digest=sha256:77154c51c768c2cfb6564a47d3074920eba6789ed361828c663eaf7d4e2a0c81

Observation efe21a7b-af83-4aea-b805-06006d46249a · outbound

This paper cites Multi-task reinforcement learning with context-based representations.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Multi-task reinforcement learning with context-based representations

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:01.310812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:53.631167Z digest=sha256:1498335591209bce6c9e84a4b56b4bab499687bbdbb672d8bebb9c27f3cd899b

Observation c9aa705f-9b4b-44d3-a5c5-4b0ef5788f1a · outbound

This paper cites T., Abdolmaleki, A., Zhang, J., Groth, O., Bloesch, M., Lampe, T., Brakel, P., Bechtle, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners T., Abdolmaleki, A., Zhang, J., Groth, O., Bloesch, M., Lampe, T., Brakel, P., Bechtle, S

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.982770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:53.780565Z digest=sha256:f0644ad768d8de6a709cdee64f2ea2ac24b90949d67464647d555d8bfc747c89

Observation 517bebf6-d433-4734-b018-4b8d1cc626b7 · outbound

This paper cites Paco: Parameter-compositional multi-task reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Paco: Parameter-compositional multi-task reinforcement learning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.682511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:53.893424Z digest=sha256:a8ecb0cdefc873293bcaf65e068974130696cb3cbce99cb50db3e1753052b5b9

Observation 3638475a-9732-4c67-a418-49cf4febf075 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.004334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.004334Z digest=sha256:1f630d8fdfa773d9861097b15ccf99f55a6be6ac8a4156fcabf832dd22648054

Observation f4797f4a-225b-42be-bab4-c83fa7fde7b1 · outbound

This paper cites DeepMind Control Suite.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DeepMind Control Suite

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.169006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.169006Z digest=sha256:5d978699eaded782c8f3d372fbd9956b969d69aa6eb13f2f88aa1b1c0aad02dd

Observation 9ad56e27-cc51-49c5-973b-9397858f3258 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Gemini: A Family of Highly Capable Multimodal Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.270164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.270164Z digest=sha256:95ab51e17201ec1ff42c84c693e52f3d8f2ac24560e5c116d4e14cd83af3a2c9

Observation 13c97a2e-7c15-4072-ada9-4c0f675e3410 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Octo: An Open-Source Generalist Robot Policy

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.408075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.408075Z digest=sha256:b2e90f4c48714632e73491cf828d5f5d8b53e5a75ace27db8a93f78b1dec6b87

Observation 4abc9372-2b55-4050-b912-923702c19438 · outbound

This paper cites M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.375399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:58:54.542081Z digest=sha256:d55b68175928bd4801124ed0f68a8a9b133bffc415c2a84a8c15c3ca8136e601

Pith citing papers

Observation bc77ead2-1f30-42fd-9916-5280d93a49ee · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:05.516882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:05.516882Z digest=sha256:17024bc63df9acf3f73ba96d8163e5b8ac1418303dbc6eeee541b016839ac37d

Observation 0b865dce-dccd-424b-8539-56b73b144627 · inbound

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning cites this paper.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 1937

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:50.060301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:50.060301Z digest=sha256:16159b1ea40ec4ebe33f4266b9a439972f41750ea6e1f3633a8f16813f266eeb

Observation 60621c76-8d7e-4e50-99fc-2dba385bb7b0 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.973880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:2f180e46828e24046c0970fc4997805363289c128fdf7a5b10ac0096c128b366

Observation 888cc10a-ff82-4837-a3e4-b20e78351f66 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:49.716390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:aacf6df7491c9fd52eb838d8768be2c6fad73b3d94dc463ff7a1881d1e6c8df5

Observation 9c5a25c9-0ec3-42b4-81d1-2f0be6bf0658 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:12:41.363803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:4ac74e53beef580b93b1cfb9f80b268d49196bbb1f81033a520ce91f903205e1

Observation 309e830e-b3ea-45f2-9ca8-588e9c61e2ee · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.607354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T05:33:19.038889Z digest=sha256:bc32f7190210a231483efd4119152d708a961e9b9ea1c56a04ed40e66f49ae74

Observation 2664d873-3fe7-4957-8bac-b87353feff58 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.804374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T06:27:38.643667Z digest=sha256:fb6b15ae1f0b46f8cd492cd74e87ef3cd7b6ba15c0d457e268769fbce2b8adbc

Observation 85285cc0-7eaf-4f1f-be3b-64a47fdd07a0 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:09:12.743772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T23:04:12.943222Z digest=sha256:2a981183c9cc09f9196b57f1d295d96904d7d61111ccde86d8f2c358bd5aad40

Observation 08c63a93-2769-4044-b758-df2313d4dca4 · inbound

Debiased Model-based Representations for Sample-efficient Continuous Control cites this paper.

Debiased Model-based Representations for Sample-efficient Continuous Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:12:28.425885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T07:09:01.132424Z digest=sha256:75694b25fed51220bf19c768da50f494e801d4e65e9675e67ca8722c1e18623b

Observation 5c937208-2081-414b-8030-b1dd12dd65b4 · inbound

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control cites this paper.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 291

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:47.358278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:47.358278Z digest=sha256:398888c19980cc94a18a207ebe1f4cb0ebf42f7749a4d812010cb7537a683b15