Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:54.542081Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 9 inbound Pith citation observations for arXiv:2505.23150.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:54.542081Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:05.516882Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T23:09:12.737900Z
100 of 118 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c403830a-7572-46f5-8a4c-4e8f99131eac · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c9da86-63a7-4a64-b95f-3a0d48e2c8c8 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Provable benefits of representational transfer in reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 713717d6-e6bd-4c9d-96bd-c26975730a9d · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Courville, A., and Bellemare, M
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a075bd-184a-4409-aee0-3305870ecba4 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Courville, A
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7da3fa4c-e85c-4342-99ec-4195f8b9d31f · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Hindsight experience replay.Advances in neural information processing systems, 30, 2017
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf984bb8-9e57-4b9d-b244-bc6220398bcc · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c0d758-659e-43e9-ab11-00f55115d6d5 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Video pretraining (vpt): Learning to act by watching unlabeled online videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98169282-9aa4-4dae-8152-7ca68d2c59fc · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Smith, L., Kostrikov, I., and Levine, S
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b1b6998-d6db-4859-ba31-625e5dc47236 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Schaul, T., van Hasselt, H
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9949fcb-109c-4545-853c-02164a0e7727 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners G., Naddaf, Y ., Veness, J., and Bowling, M
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f64815-4508-4e5a-ac03-e3cbf631d284 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners G., Dabney, W., and Munos, R
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08d1053d-46cb-4f71-880d-78bc3dd62be7 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Dynamic Programming
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2486c5fb-c6b4-443c-8cdf-fd42dd633eb4 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93b4fe42-9f73-4d76-82fa-1430bf502f95 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners P., and Weinberger, K
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d41d36-8aae-423b-996d-11f1b1fdfef7 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c3a67c-9aca-4c18-9056-0c246c26dc54 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners RT-1: Robotics Transformer for Real-World Control at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 778f8cbc-d862-4bf3-b45f-d40d66689849 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8a57e69-4a22-4384-9865-e7bbadf3fcad · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Decision transformer: Reinforcement learning via sequence modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bfdf1d4-ebbf-4fa2-969a-3b77060d52cf · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A system for general in-hand object re-orientation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c71d3a6-7cbf-4467-85f6-f06c0eb2f829 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Gradnorm: Gradient normaliza- tion for adaptive loss balancing in deep multitask networks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b133e83c-fc66-447e-8f32-286d8c5f2621 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Just pick a sign: Optimizing deep multitask models with gradient sign dropout
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 725d0859-c737-4840-9e5f-ba3f83d84306 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5752454-cbfc-4efa-9bae-76e8fa6c8afe · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f173b2ca-f1b5-45e1-a98b-0ac0ff439627 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Better exploration with optimistic actor critic
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9579dbe-bdc1-4982-8b39-ed18dca7ba0f · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Magnetic control of tokamak plasmas through deep reinforcement learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 544c6b33-a6c6-45fd-ac93-60cd1d1f1f1d · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cedc50c4-c18d-4625-9599-87af9b1a4611 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners An image is worth 16x16 words: Transformers for image recognition at scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e88b3017-0087-4ffc-864a-353bd7bbd744 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8156fac5-839e-4306-9452-eb846bba34f8 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a990062a-6dba-4f00-b38e-c2f4518b63d4 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A., Chebotar, Y ., Xiao, T., Irpan, A., Levine, S., Castro, P
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba3cede9-4995-41b8-a536-f8b58699f1f3 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Model-agnostic meta-learning for fast adaptation of deep networks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e78924f-1e27-4136-9ecb-d46590b3ef89 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Gu, S
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f8af50-5032-4c35-8ef3-b0e167b5c22f · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Addressing function approximation error in actor-critic methods
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad6cab91-fa5f-41a7-bd2b-995743cf13f6 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Precup, D., and Meger, D
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5c7ad6-7ec2-46ff-b60d-2c762b03bdd6 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Divide-and-conquer reinforcement learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11695065-eb2b-43bf-8acd-b1b64740f6bb · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b11e6db-448c-4902-ad8b-699cd749dc7e · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Mastering Diverse Domains through World Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd824f7-6396-4ee8-87f5-7728212e1b21 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners TD-MPC2: Scalable, Robust World Models for Continuous Control
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee899869-f4ac-402a-bf1a-243769f12730 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners R., Millman, K
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83679e9d-c067-4d62-96aa-c79d0a711e38 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners T., Wang, Z., Heess, N., and Riedmiller, M
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9bb63e-08e0-494c-bd15-61e1e3d3a57a · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient multi-task reinforcement learning with cross-task policy guidance
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e672c4-b654-4d62-9126-c99e3324d0bc · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling Laws for Autoregressive Generative Modeling
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa26013b-9414-44dd-8134-fb1b1af1cdd2 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Multi-task deep reinforcement learning with popart
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10abf86-0652-4874-9896-dc073ebdf597 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Training Compute-Optimal Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da2ff812-3cf0-4a19-a9f0-32166cf81636 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Otter: A vision-language-action model with text-aware visual feature extraction
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d70b86f-53ea-46b7-86c0-9f050d402e4c · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 433abc24-6485-4cf0-9535-847bfb6ac080 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c78482a-fe9b-43f3-9780-a093bedcb51a · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a618c6a6-b9fc-42b8-81af-bfefb5440eaa · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling Laws for Neural Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2632db-b1fc-431f-b978-3e0be2c69394 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Champion-level drone racing using deep reinforcement learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9dd9a88f-9958-4e08-a061-06a00151d3e1 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners OpenVLA: An Open-Source Vision-Language-Action Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ba696b-a068-4fdc-942f-828e6bd8b89a · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Ba, J
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25391644-5245-47ae-a423-bd5dfe19e699 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners C., Lo, W.-Y ., et al
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dabe2b8f-9758-4483-88a5-a239cd795e92 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b377d3d-6be3-4aa6-b19d-e7f2245a7783 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0fce54b-56f3-4e46-9d78-a750c2e7257d · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92fdb5d-b34e-42b9-9dd8-0b79474c4011 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners RMA: Rapid Motor Adaptation for Legged Robots
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b884ea-e9b5-4d08-8c9d-2d6a9dafe2a9 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Offline q-learning on diverse multi-task data both scales and generalizes
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e53f8be0-5c13-4972-8dd5-fe5ce4b48e33 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Reinforcement learning with augmented data
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e598f6e-02ff-41e2-80f0-7de53dca1f64 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac955de-f366-46dd-950e-f7f7e1d1f34f · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Hyperspherical Normalization for Scalable Deep Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae14b0df-74e1-4d93-bc6c-6e594c8ee1ac · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners End-to-end training of deep visuomotor policies
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba682875-3697-4628-a572-3f871e74de31 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DeepSeek-V3 Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 391030fa-4064-41f5-a09d-32a19b9a97b0 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Conflict-averse gradient descent for multi-task learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff021dd4-6f61-420c-ad80-4985741505ef · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91825c01-af53-4a35-827c-03fa771c2fe6 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling laws for fine-grained mixture of experts
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5446eea-762d-424c-bc91-e942572d7a78 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a72955a-105a-4fd6-824e-1a6f81559308 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners An Empirical Model of Large-Batch Training
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0add11d2-7c60-4e76-9f95-c193c5bd944f · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A., Veness, J., Bellemare, M
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5a37fcc-0fb6-4ddd-b758-ab4199734896 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abafc6a0-52de-4d8d-a1fa-c3a640fc15c7 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Tactical optimism and pessimism for deep reinforcement learning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f320d94-7ca4-4ed6-b067-77c05124fbfe · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Cygan, M
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19548d37-4bec-43bc-af16-66f226479d82 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b79eba20-6d7e-4ddb-b599-54112b534780 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc34fa6c-1f32-40c8-b836-35a24ae65bc6 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Dinov2: Learning robust visual features without supervision
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08140eb0-776a-4c7a-be59-1307dba4a575 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Training language models to follow instructions with human feedback
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bba43ec-1eaa-4744-9ece-92d7769e9d23 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4eaee66-7207-4e20-a551-212ec8395e2c · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Is Value Learning Really the Main Bottleneck in Offline RL?
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4dbe154-c0a7-4194-aec4-404adbed9de7 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Flow Q-Learning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b01f985-6863-4184-b4fd-d48391dbff62 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9ea69e-97fa-4277-973e-b05b52d26d1f · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f3c6d5b-3fbd-431d-b5ed-0507340cfac5 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient off-policy meta- reinforcement learning via probabilistic context variables
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0935f476-d90e-4a90-ab89-f9845ff239e6 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A Generalist Agent
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe4849cb-8398-425d-b46a-d6a609b6f0a4 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Distillation
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c54498a5-69f8-4676-9eeb-e9dc48e64daa · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Value-Based Deep RL Scales Predictably
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312c1e26-ea4e-4218-b376-37e15e0b2dfc · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners D., Courville, A., and Bachman, P
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bf15b1f-6df2-48bb-9ed5-19c29a58c920 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 980db95b-203d-45db-9dc7-f0f6d759858c · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Koltun, V
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26518396-25da-4193-9b3b-c8c1a5ad8717 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d99e650-1b44-4599-ae47-53f38f3ab586 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d63ad454-a067-4c4b-8a9c-4c038b1d35fd · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Deterministic policy gradient algorithms
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6390f0ca-ea85-4fef-9b4a-b50d88db26b9 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Mastering the game of go without human knowledge
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efe21a7b-af83-4aea-b805-06006d46249a · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Multi-task reinforcement learning with context-based representations
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9aa705f-9b4b-44d3-a5c5-4b0ef5788f1a · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners T., Abdolmaleki, A., Zhang, J., Groth, O., Bloesch, M., Lampe, T., Brakel, P., Bechtle, S
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 517bebf6-d433-4734-b018-4b8d1cc626b7 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Paco: Parameter-compositional multi-task reinforcement learning
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3638475a-9732-4c67-a418-49cf4febf075 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4797f4a-225b-42be-bab4-c83fa7fde7b1 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DeepMind Control Suite
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad56e27-cc51-49c5-973b-9397858f3258 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Gemini: A Family of Highly Capable Multimodal Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13c97a2e-7c15-4072-ada9-4c0f675e3410 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Octo: An Open-Source Generalist Robot Policy
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4abc9372-2b55-4050-b912-923702c19438 · outbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc77ead2-1f30-42fd-9916-5280d93a49ee · inbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b865dce-dccd-424b-8539-56b73b144627 · inbound
RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 1937
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60621c76-8d7e-4e50-99fc-2dba385bb7b0 · inbound
What Does Flow Matching Bring To TD Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 888cc10a-ff82-4837-a3e4-b20e78351f66 · inbound
FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c5a25c9-0ec3-42b4-81d1-2f0be6bf0658 · inbound
FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 309e830e-b3ea-45f2-9ca8-588e9c61e2ee · inbound
When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2664d873-3fe7-4957-8bac-b87353feff58 · inbound
When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85285cc0-7eaf-4f1f-be3b-64a47fdd07a0 · inbound
When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08c63a93-2769-4044-b758-df2313d4dca4 · inbound
Debiased Model-based Representations for Sample-efficient Continuous Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.