Pith. sign in

Paper Citation Record · LEDGER

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM

As of 20 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2505.10861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10861 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:05:21.477420Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:25:37.918588Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation 500b0b05-6bbd-4c6d-b30e-4e3099c7708f · outbound

This paper cites Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.202033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.202033Z digest=sha256:c8cb0e0b8c521662a94f7230c4873682159813f6cf16f6219dcb5d41a3d2dab4

Observation 945fb16b-7968-456a-8893-509d74674454 · outbound

This paper cites Reinforcement learning: Theory and algorithms.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Reinforcement learning: Theory and algorithms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.208253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.208253Z digest=sha256:c1df16eb862e21a338c8ffa0b4d3c686e20e1e3cffe5f4d220c939d7c69134ee

Observation f975919c-8866-4ea1-8313-1a83144506b7 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.213425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.213425Z digest=sha256:856cb332864f829375a9e73b35c8aece7f44da9d68e8f6f337d1222c791064b1

Observation 807bbdff-39eb-47b5-aca6-892ea3e9b929 · outbound

This paper cites Unsupervised State Representation Learning in Atari.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Unsupervised State Representation Learning in Atari

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.218302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.218302Z digest=sha256:60193104c167d5b04d00bb406f1b967e8e835c19fc027c1f06bb6459c7caa545

Observation e7b764f1-0e47-4bab-bbae-5429a9dd66e3 · outbound

This paper cites Griffiths.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Griffiths

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.224310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.224310Z digest=sha256:6854e157bea9a8bf480578c59787c1bd4dc537119c4f3c5e3693b9a787f9efd7

Observation d6c5c746-9aae-4a49-8511-fc163fa3c56f · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.229109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.229109Z digest=sha256:4e7656d5c6071df876a1e71455ebc52058271eccc9c73b31536c6f3240d8df2e

Observation 1e92dbc5-5abb-4937-9fbe-2a8046b82339 · outbound

This paper cites Neuro-dynamic programming: An overview and recent results.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Neuro-dynamic programming: An overview and recent results

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.388449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.234720Z digest=sha256:ea026cde070ba2e53700c46d48733c580a43d5acf248e9335f4ce38b4e444e26

Observation aecc0872-0d13-4558-85bb-7a2e79c1f3e8 · outbound

This paper cites Grounding llms for robot task planning using closed-loop state feedback.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Grounding llms for robot task planning using closed-loop state feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.239027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.239027Z digest=sha256:28071e46dfc29de24efcfac6b12be9f883734700cea3702fe13f29da7d3bd263

Observation 7012707f-78d1-4f4a-a0c8-7bf0b23ccd7f · outbound

This paper cites Language models are few-shot learners.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Language models are few-shot learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.243611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.243611Z digest=sha256:aef5961a7f00df5f14cf53a150c7ebbfa7594e5e319a550b833dd46084834ea2

Observation 803e20cd-edfe-4994-ba6e-0f4f88ca2145 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Grounding large language models in interactive environments with online reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.365281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.248737Z digest=sha256:5f8416ccd4791f9637915f7ef07bfc8ba7c969a425d5945a9503799657cfe7c4

Observation 6af3f052-b3ce-4e7d-bc9b-6e9753567a0c · outbound

This paper cites Efficient Sequential Decision Making with Large Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Efficient Sequential Decision Making with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.253125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.253125Z digest=sha256:99aa7dd9cd53a30daf65a5659d86ca0d6745b8612a2c832d8bc1630ac0c31cd3

Observation 8efc4fdf-dced-47c6-98b7-1a849c8ff17c · outbound

This paper cites LMPriors: Pre-Trained Language Models as Task-Specific Priors.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM LMPriors: Pre-Trained Language Models as Task-Specific Priors

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.257799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.257799Z digest=sha256:33759a796ddeba27216248c138ded13c441f1915c032d0d96776aa2a9e4a551e

Observation 74d1beda-5d17-4ad2-8bc7-b22490e7fa56 · outbound

This paper cites Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.262506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.262506Z digest=sha256:29a961e59d6aa259a4c34a287e16571214712898d5649bb67f6f64e7a12b96cc

Observation 086f8a74-0640-4fcc-b6fb-0ee254267183 · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Guiding pretraining in reinforcement learning with large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.267546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.267546Z digest=sha256:a248c8684e2350950e1f6713a227bb179058aef146d6b3ebf9ce4d130070e5f2

Observation 04171026-0468-4cc0-952c-e7e85607e395 · outbound

This paper cites Tree-based batch mode reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Tree-based batch mode reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.272038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.272038Z digest=sha256:8f07dd172a67e0e95ee222ea0a823649fda59ec138ade9d17d8f67dce75a6e94

Observation a48a4f33-20fb-4817-8add-79a50cc727b9 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Soft Actor-Critic Algorithms and Applications

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.276339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.276339Z digest=sha256:fac004d977076753506f6d2167d1050be7023e6346e053ac7383e60abfd8c60b

Observation faa31b6e-e34a-4756-9890-4308bb567c7e · outbound

This paper cites Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.281529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.281529Z digest=sha256:1edc5101dab64c03ff0bbbdbeaa0b00bf51d03f0e3a0b73522624c6c0294e918

Observation e564af48-2bb2-4c83-a939-f96492fa3324 · outbound

This paper cites Deep q-learning from demonstrations.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep q-learning from demonstrations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.286683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.286683Z digest=sha256:eed66c0d61d32538fcc2199fd989b1aa17529d388ce1a4d462d9a83c09d9d6ef

Observation 8ca3dfa6-a5f3-482e-a43a-1373fbd0dd2c · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM 3d-llm: Injecting the 3d world into large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.291172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.291172Z digest=sha256:ad4541eec56423764c64eaa1e15d06be3de5bc76dbe94efa3853a07ec219ad8d

Observation 6f6fa83f-5aa2-4d06-8703-9c123f6831af · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.295393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.295393Z digest=sha256:5d65931c076b4ef1a27e621341fdec6a556c718aea2edefa424a4d1d8edeed03

Observation d578f68d-f64b-4130-8c0b-34a273fc3554 · outbound

This paper cites Visual language maps for robot navigation.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Visual language maps for robot navigation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.314781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.299703Z digest=sha256:6cafdc597153883f41c796af26e4403d12b323bcc2421c13201d9b240226dbb5

Observation adfed43b-9118-4cf9-b9ca-0b365f91c080 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.304138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.304138Z digest=sha256:93c56757a53dc71a332c6bee95212bd922d2456a490545b9b6bfee5bf1edcb44

Observation 97f35f64-43b2-4862-b800-030416b775f9 · outbound

This paper cites A survey of robot intelligence with large language models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM A survey of robot intelligence with large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.308472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.308472Z digest=sha256:0b851c492b25301c5fc8d24206ceaad3341a5dc7fb8f397b0d79ea653f8c04d1

Observation 9ade9eef-bd05-4d06-82a1-7e3fd16b17a6 · outbound

This paper cites Bench llm deciders with gym translators.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Bench llm deciders with gym translators

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.300138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.313022Z digest=sha256:6cf67b329ae1af3fb94954cdd8113ad69c17857a2e77aa0ab45f3b3b9b4fac9a

Observation aa5fcee6-4bf2-4b9f-a749-6241563b3065 · outbound

This paper cites Housekeep: Tidying virtual households using commonsense reasoning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Housekeep: Tidying virtual households using commonsense reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.285482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.317283Z digest=sha256:8a0170afe33291459f80537d8e17042128a2698877d8efd93d7b33d8539aa982

Observation a3311083-fba8-45a7-b497-08972e02ca09 · outbound

This paper cites Reinforcement learning in robotics: A survey.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Reinforcement learning in robotics: A survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.321708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.321708Z digest=sha256:272c90e3b1c15f49c5b00415563334d4115d32d465690ccd7370f59c35d78b2a

Observation 0fa4467a-407f-4a72-b660-81f06ff2e7fc · outbound

This paper cites Can large language models explore in-context?.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Can large language models explore in-context?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.325869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.325869Z digest=sha256:f4b2d030935758725a1319e0865250b6257f499cb38db5d9557185d673451cac

Observation d236be72-1f30-400a-b6f9-fc799da41d2f · outbound

This paper cites Batch reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Batch reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.330290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.330290Z digest=sha256:8883ec54e7b34ccfc3a0f2710d26196ea9439e18a6ba00c406259d3d6ca39d39

Observation 41565e31-b181-4d8e-9d97-25c573350a16 · outbound

This paper cites Supervised pretraining can learn in-context reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Supervised pretraining can learn in-context reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.335032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.335032Z digest=sha256:6ae198c309f3998a5ce32ec044152d38308e31381b1573de81b05c8570b3eca4

Observation 6bc88d04-51be-41ef-98bd-feb84b348992 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.339425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.339425Z digest=sha256:239c934dc102a0917ce1ea5f4133062b1d922f62c8135258b52a535694ede8f8

Observation 1ac0219f-c239-4f1d-b07b-1b8156ccb9f4 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Code as policies: Language model programs for embodied control

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.344099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.344099Z digest=sha256:dc86ea49a00e5fa946382361cea29aff2a5199283e7e34ed691cdde248e9a1f7

Observation a90147ac-36aa-4152-be0b-12e7a0655000 · outbound

This paper cites Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.348912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.348912Z digest=sha256:777cccbddde3c40b3e16a78616af3f7539d976f7fbf2363e2cd38d65453a5820

Observation 07649708-e6e6-45e2-82aa-4f5d76599e6d · outbound

This paper cites Physgen: Rigid-body physics-grounded image-to-video generation.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Physgen: Rigid-body physics-grounded image-to-video generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.236332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.353765Z digest=sha256:ac22253199b4f8f934c70e7a10046786ccd8a080617b5a5ed62af93cf2d2202e

Observation 576cefc7-e47d-47c8-aff9-632f26c4d459 · outbound

This paper cites A Survey of Reinforcement Learning Informed by Natural Language.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM A Survey of Reinforcement Learning Informed by Natural Language

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.358359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.358359Z digest=sha256:b13f09d4051b1a5a26173b3eaec30e31f892b1bbdebf205d63908c12b7debdc5

Observation be556654-9cbb-4619-8a4f-accd09f8c760 · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.363027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.363027Z digest=sha256:14ab87bcaab36e8f9b9935ff3142887afb8dbf392acf3b9895ea3a5052083775

Observation 2db7ad14-22a4-48dc-bccd-84606c27c9b8 · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Overcoming exploration in reinforcement learning with demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.367902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.367902Z digest=sha256:b591f1a6bdcf1f75fa03f3bffe018effb6657257f9e943fff7c098bce84ef61b

Observation 52fc9b45-c769-42c5-8c02-f35e96739d8b · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.372358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.372358Z digest=sha256:f2f86c1a1ca3d7f0b2d89336266be8bbc5f0f020995f550d6a4af7e5b50a1ce2

Observation e94f63ce-d0b5-400a-a8e0-1f0595fdf8e8 · outbound

This paper cites EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.377077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.377077Z digest=sha256:3e6f95d29d75efe2ddf1ec67b668474e4bec537779a64e4ba7c28395e32ca48a

Observation d095b5cb-644b-4028-8449-86f33522ddb5 · outbound

This paper cites Llamagym: Fine-tune llm agents with online reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Llamagym: Fine-tune llm agents with online reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.213175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.381607Z digest=sha256:1c3accaf028a4b52d91c291b9f98c6443b3ed7dc62deef2ba51d7a898c92fa1d

Observation 6db3aedd-05f8-478f-af60-1ef3638e7622 · outbound

This paper cites Mapping language models to grounded conceptual spaces.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Mapping language models to grounded conceptual spaces

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.385866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.385866Z digest=sha256:5a1c7f5e59c8d23c487624ec1da3bd6259f15f9be254290b08c94d97e73f1e0a

Observation e3d43e25-5807-4d58-a31f-eca6ac07c11c · outbound

This paper cites Deep Q-Learning: Theoretical Insights from an Asymptotic Analysis.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep Q-Learning: Theoretical Insights from an Asymptotic Analysis

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:05:21.695486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.390484Z digest=sha256:c9506d753b090195d05bdb63e6e990b2228be5f214df77b29019d6e04084d657

Observation 92b22de6-bf68-4b0e-bc43-45cd26ac1e31 · outbound

This paper cites Neural fitted q iteration--first experiences with a data efficient neural reinforcement learning method.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Neural fitted q iteration--first experiences with a data efficient neural reinforcement learning method

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.187685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.395272Z digest=sha256:baaf591533f327cacde7bdd6501790cfe39017d21627977749a3b4e8b0ebfd26

Observation f695cd43-3bd0-4c65-80d6-19a3efb495fb · outbound

This paper cites Kickstarting Deep Reinforcement Learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Kickstarting Deep Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.399556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.399556Z digest=sha256:635085087db92758036aefcebd4be30b3cfdfaf822d47d2843f9e025449b88a1

Observation dd8f0dfc-628e-46a6-a737-0f46f2607295 · outbound

This paper cites d3rlpy: An offline deep reinforcement learning library.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM d3rlpy: An offline deep reinforcement learning library

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.404084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.404084Z digest=sha256:1010e510328513d9f42ca70f077318b6f41344b1f1ef254b2141e3fd37aea50d

Observation b0e751dd-a5ec-47a2-ac7f-a1c8ee1cecf4 · outbound

This paper cites Mastering the game of go without human knowledge.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Mastering the game of go without human knowledge

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.408938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.408938Z digest=sha256:cc38508b16898a0d6c287a6a1b963236fcfbf50891a2c3e65465b4eafe681f59

Observation 97cc27ab-56ac-4efe-8d15-7e70b2245994 · outbound

This paper cites It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.413563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.413563Z digest=sha256:cf685a009643d987087e4ee595170958c74e00cb530c0b6d6586cc5421ae2b94

Observation 143dc137-5a88-49bb-ae91-a05c169d57b4 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.418269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.418269Z digest=sha256:1477f47970bf7d2cf71467c5215eb9ea3ed2d5282da31fba5c8bb1b6cc58c670

Observation 8a37a8d8-fd26-42ef-aa7e-625726d1c3a1 · outbound

This paper cites True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.423074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.423074Z digest=sha256:db5d3be4478c1c08bcb08eaea7f029eade651715e340902d2186869b23d5e7cb

Observation 3e024e96-912d-454a-80e5-b4b9abfacfc5 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.427700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.427700Z digest=sha256:2af1169dea7671055cd0dba67d9ddc8062c2527124c902e5fa1553b4006f8c4f

Observation 6dcd2c35-d5d6-4a16-ad07-9515d53786e3 · outbound

This paper cites Representation Learning for Online and Offline RL in Low-rank MDPs.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Representation Learning for Online and Offline RL in Low-rank MDPs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.432391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.432391Z digest=sha256:d6daf0956ed649a97d307540cf6dc479c6e425f1090a00126683ebc5cc76b253

Observation a59c1136-1d1b-4ce4-8d72-2b6a99b53dbd · outbound

This paper cites Deep Reinforcement Learning with Double Q-learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep Reinforcement Learning with Double Q-learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.437109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.437109Z digest=sha256:a20b2d5b84909d58b3ee9c2d93eb7a45f778cdc2c4ac6ca0a886d8fd3d6ebe13

Observation 0d28139d-0e1b-4166-b820-945784beb028 · outbound

This paper cites Instabilities of offline rl with pre-trained neural representation.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Instabilities of offline rl with pre-trained neural representation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.153897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.441831Z digest=sha256:a3c0a78961c00cc5ed5b653f050aa4c3d800c35a41b1ef444ec2739462699810

Observation 1b2f8424-0178-4a98-a6f5-b7161f2540a6 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Chain-of-thought prompting elicits reasoning in large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.445957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.445957Z digest=sha256:030c8d3d0b638529f5791f4df00f7a50ddd26f2e793cf691ee496ce1283e6876

Observation c4e53891-9749-4d06-a701-4061e7827b72 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.129447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.450587Z digest=sha256:b3759fca0ee48b678234e1f2959d61167abc151c761a1de7be9c5b0c974912cc

Observation a61e155c-1586-4e41-af76-13ed7df7292f · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.454824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.454824Z digest=sha256:a70e8581c6dadcd82ba89bf073a1807a922e9c946d6961f3d4091643f142ccb3

Observation bd68ada6-c12b-441e-8fe8-4b81eb54573e · outbound

This paper cites Qwen2.5 Technical Report.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Qwen2.5 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.459606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.459606Z digest=sha256:5e0f8564654ae6224ca8133455b9aaa45562ac0c1dba529ce793052fdd8ec61d

Observation 9b2129aa-1500-4fb9-b8a8-401021eb3933 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM ReAct: Synergizing Reasoning and Acting in Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.463838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.463838Z digest=sha256:73661eacc4fa598a89b49a7d2add01ff2a7cd5a0fa8af3ac55069d66c0668e03

Observation 41044195-3ce1-4045-a71f-decbeb0082d3 · outbound

This paper cites Policy finetuning in reinforcement learning via design of experiments using offline data.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Policy finetuning in reinforcement learning via design of experiments using offline data

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.114177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.468622Z digest=sha256:0c8a6799606c05d38243a41e8a8d384c0cc2bfcd10a26094bdfa288ffae23a2d

Observation a4504fb6-8505-49b8-b05a-66dc89d85d30 · outbound

This paper cites Online decision transformer.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Online decision transformer

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.097815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.472787Z digest=sha256:d26e808da44060455f61a4c150b801598605a752acac9e207b83890b182e9460

Observation 2e7e81b9-3538-4909-9f76-5aa26cd04b72 · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.477420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.477420Z digest=sha256:4082667da56a59f8291a9615eb0a0a24c16f6a0bacb13316e1e1b96a089e7128

Pith citing papers

Observation f2d30702-5654-451e-a9d5-872b374fdef9 · inbound

ProDVI: Programmatic Dynamics Priors for Value Network Initialization cites this paper.

ProDVI: Programmatic Dynamics Priors for Value Network Initialization Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:25:38.695370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:25:37.918588Z digest=sha256:1e91ed804bb5636c4d9bc8d7ea54787e11443da9b55415010ba2fa4c1cde81b1