Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:05:21.477420Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2505.10861.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:05:21.477420Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T19:25:37.918588Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T05:30:23.456663Z
60 of 60 outbound references displayed
External citation measurements
0
pith, observed 2026-08-10T05:30:23.456663Z
Observation 500b0b05-6bbd-4c6d-b30e-4e3099c7708f · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 945fb16b-7968-456a-8893-509d74674454 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Reinforcement learning: Theory and algorithms
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f975919c-8866-4ea1-8313-1a83144506b7 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 807bbdff-39eb-47b5-aca6-892ea3e9b929 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Unsupervised State Representation Learning in Atari
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b764f1-0e47-4bab-bbae-5429a9dd66e3 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Griffiths
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c5c746-9aae-4a49-8511-fc163fa3c56f · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Efficient Online Reinforcement Learning with Offline Data
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e92dbc5-5abb-4937-9fbe-2a8046b82339 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Neuro-dynamic programming: An overview and recent results
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aecc0872-0d13-4558-85bb-7a2e79c1f3e8 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Grounding llms for robot task planning using closed-loop state feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7012707f-78d1-4f4a-a0c8-7bf0b23ccd7f · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Language models are few-shot learners
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803e20cd-edfe-4994-ba6e-0f4f88ca2145 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Grounding large language models in interactive environments with online reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6af3f052-b3ce-4e7d-bc9b-6e9753567a0c · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Efficient Sequential Decision Making with Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8efc4fdf-dced-47c6-98b7-1a849c8ff17c · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM LMPriors: Pre-Trained Language Models as Task-Specific Priors
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d1beda-5d17-4ad2-8bc7-b22490e7fa56 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 086f8a74-0640-4fcc-b6fb-0ee254267183 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Guiding pretraining in reinforcement learning with large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04171026-0468-4cc0-952c-e7e85607e395 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Tree-based batch mode reinforcement learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48a4f33-20fb-4817-8add-79a50cc727b9 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Soft Actor-Critic Algorithms and Applications
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa31b6e-e34a-4756-9890-4308bb567c7e · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e564af48-2bb2-4c83-a939-f96492fa3324 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep q-learning from demonstrations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ca3dfa6-a5f3-482e-a43a-1373fbd0dd2c · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM 3d-llm: Injecting the 3d world into large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6fa83f-5aa2-4d06-8703-9c123f6831af · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM LoRA: Low-Rank Adaptation of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d578f68d-f64b-4130-8c0b-34a273fc3554 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Visual language maps for robot navigation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation adfed43b-9118-4cf9-b9ca-0b365f91c080 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Inner Monologue: Embodied Reasoning through Planning with Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f35f64-43b2-4862-b800-030416b775f9 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM A survey of robot intelligence with large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ade9eef-bd05-4d06-82a1-7e3fd16b17a6 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Bench llm deciders with gym translators
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aa5fcee6-4bf2-4b9f-a749-6241563b3065 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Housekeep: Tidying virtual households using commonsense reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a3311083-fba8-45a7-b497-08972e02ca09 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Reinforcement learning in robotics: A survey
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa4467a-407f-4a72-b660-81f06ff2e7fc · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Can large language models explore in-context?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d236be72-1f30-400a-b6f9-fc799da41d2f · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Batch reinforcement learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41565e31-b181-4d8e-9d97-25c573350a16 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Supervised pretraining can learn in-context reinforcement learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bc88d04-51be-41ef-98bd-feb84b348992 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac0219f-c239-4f1d-b07b-1b8156ccb9f4 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Code as policies: Language model programs for embodied control
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90147ac-36aa-4152-be0b-12e7a0655000 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07649708-e6e6-45e2-82aa-4f5d76599e6d · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Physgen: Rigid-body physics-grounded image-to-video generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 576cefc7-e47d-47c8-aff9-632f26c4d459 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM A Survey of Reinforcement Learning Informed by Natural Language
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be556654-9cbb-4619-8a4f-accd09f8c760 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Eureka: Human-Level Reward Design via Coding Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db7ad14-22a4-48dc-bccd-84606c27c9b8 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Overcoming exploration in reinforcement learning with demonstrations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52fc9b45-c769-42c5-8c02-f35e96739d8b · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e94f63ce-d0b5-400a-a8e0-1f0595fdf8e8 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d095b5cb-644b-4028-8449-86f33522ddb5 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Llamagym: Fine-tune llm agents with online reinforcement learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6db3aedd-05f8-478f-af60-1ef3638e7622 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Mapping language models to grounded conceptual spaces
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d43e25-5807-4d58-a31f-eca6ac07c11c · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep Q-Learning: Theoretical Insights from an Asymptotic Analysis
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 92b22de6-bf68-4b0e-bc43-45cd26ac1e31 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Neural fitted q iteration--first experiences with a data efficient neural reinforcement learning method
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f695cd43-3bd0-4c65-80d6-19a3efb495fb · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Kickstarting Deep Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8f0dfc-628e-46a6-a737-0f46f2607295 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM d3rlpy: An offline deep reinforcement learning library
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e751dd-a5ec-47a2-ac7f-a1c8ee1cecf4 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Mastering the game of go without human knowledge
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97cc27ab-56ac-4efe-8d15-7e70b2245994 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143dc137-5a88-49bb-ae91-a05c169d57b4 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a37a8d8-fd26-42ef-aa7e-625726d1c3a1 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e024e96-912d-454a-80e5-b4b9abfacfc5 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dcd2c35-d5d6-4a16-ad07-9515d53786e3 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Representation Learning for Online and Offline RL in Low-rank MDPs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a59c1136-1d1b-4ce4-8d72-2b6a99b53dbd · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep Reinforcement Learning with Double Q-learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d28139d-0e1b-4166-b820-945784beb028 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Instabilities of offline rl with pre-trained neural representation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1b2f8424-0178-4a98-a6f5-b7161f2540a6 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Chain-of-thought prompting elicits reasoning in large language models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4e53891-9749-4d06-a701-4061e7827b72 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a61e155c-1586-4e41-af76-13ed7df7292f · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd68ada6-c12b-441e-8fe8-4b81eb54573e · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Qwen2.5 Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2129aa-1500-4fb9-b8a8-401021eb3933 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM ReAct: Synergizing Reasoning and Acting in Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41044195-3ce1-4045-a71f-decbeb0082d3 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Policy finetuning in reinforcement learning via design of experiments using offline data
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a4504fb6-8505-49b8-b05a-66dc89d85d30 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Online decision transformer
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2e7e81b9-3538-4909-9f76-5aa26cd04b72 · outbound
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d30702-5654-451e-a9d5-872b374fdef9 · inbound
ProDVI: Programmatic Dynamics Priors for Value Network Initialization Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.