Pith. sign in

Paper Citation Record · LEDGER

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM

As of 16 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2505.10861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10861 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:05:21.477420Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:25:37.918588Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation 500b0b05-6bbd-4c6d-b30e-4e3099c7708f · outbound

This paper cites Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.202033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.202033Z digest=sha256:7207f90e86cfe7d0366fb07d05fe66fbea9e90b4202d115e25b0566599b2c55e

Observation 945fb16b-7968-456a-8893-509d74674454 · outbound

This paper cites Reinforcement learning: Theory and algorithms.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Reinforcement learning: Theory and algorithms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.208253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.208253Z digest=sha256:bd6ef6af6439602361b181f398ee0df74110ab06807a472bf593e699cf6bc562

Observation f975919c-8866-4ea1-8313-1a83144506b7 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.213425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.213425Z digest=sha256:42c0b1e83c6141a6b13a4888956ef4f88eb9d0aabfa0f1702b07cd325594db5c

Observation 807bbdff-39eb-47b5-aca6-892ea3e9b929 · outbound

This paper cites Unsupervised State Representation Learning in Atari.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Unsupervised State Representation Learning in Atari

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.218302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.218302Z digest=sha256:087e3ff7d903c1c658638e6ce1a46e31e90d47639e2d5134ccd9050cc339c43f

Observation e7b764f1-0e47-4bab-bbae-5429a9dd66e3 · outbound

This paper cites Griffiths.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Griffiths

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.224310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.224310Z digest=sha256:7fed862fdba6543216b4d56c937a1f3659f0fd938076480706f206d610ef613d

Observation d6c5c746-9aae-4a49-8511-fc163fa3c56f · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.229109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.229109Z digest=sha256:834626094c6cff1967393c7bc7af070f7e9baf49db67c464ad9d38de5f0abf39

Observation 1e92dbc5-5abb-4937-9fbe-2a8046b82339 · outbound

This paper cites Neuro-dynamic programming: An overview and recent results.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Neuro-dynamic programming: An overview and recent results

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.388449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.234720Z digest=sha256:c81d336fc5add14531fef817b9ec374a94f8cecacc2679d2870b06455f225608

Observation aecc0872-0d13-4558-85bb-7a2e79c1f3e8 · outbound

This paper cites Grounding llms for robot task planning using closed-loop state feedback.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Grounding llms for robot task planning using closed-loop state feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.239027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.239027Z digest=sha256:afcabf96f01a84aac61a3a78d1060a3310f32127361d0880976b564694610ac8

Observation 7012707f-78d1-4f4a-a0c8-7bf0b23ccd7f · outbound

This paper cites Language models are few-shot learners.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Language models are few-shot learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.243611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.243611Z digest=sha256:8690c7264e655944c8a3de2fc84c2c79f2a5507d0ed652b681064055f8d42b28

Observation 803e20cd-edfe-4994-ba6e-0f4f88ca2145 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Grounding large language models in interactive environments with online reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.365281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.248737Z digest=sha256:99860d9dcf7c5fd416fa854021274f5e95112816c725526ed562d5742d6bdf0d

Observation 6af3f052-b3ce-4e7d-bc9b-6e9753567a0c · outbound

This paper cites Efficient Sequential Decision Making with Large Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Efficient Sequential Decision Making with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.253125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.253125Z digest=sha256:7ffa1b2aaaf712e6deddd1abe7a58143ea16c10d3480862880b06405d3496213

Observation 8efc4fdf-dced-47c6-98b7-1a849c8ff17c · outbound

This paper cites LMPriors: Pre-Trained Language Models as Task-Specific Priors.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM LMPriors: Pre-Trained Language Models as Task-Specific Priors

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.257799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.257799Z digest=sha256:027daedacbad5d666cf62840848ca3087d71347f8322f36445108ec0d4af9112

Observation 74d1beda-5d17-4ad2-8bc7-b22490e7fa56 · outbound

This paper cites Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.262506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.262506Z digest=sha256:9b9b1ace3a344cf13ff1aedfa8d1a29314fc002189ea46063d48bc7054191f07

Observation 086f8a74-0640-4fcc-b6fb-0ee254267183 · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Guiding pretraining in reinforcement learning with large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.267546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.267546Z digest=sha256:cf965d92121e911478b71e5d4ed625d99f3bae56221d1c5a5de434126ef9d197

Observation 04171026-0468-4cc0-952c-e7e85607e395 · outbound

This paper cites Tree-based batch mode reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Tree-based batch mode reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.272038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.272038Z digest=sha256:9222545f10ac83fded45f232b246c15ff0bc98821e86b810731699f170a85231

Observation a48a4f33-20fb-4817-8add-79a50cc727b9 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Soft Actor-Critic Algorithms and Applications

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.276339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.276339Z digest=sha256:4482240385be8b4bb06467014b760695e2adf5503b171c915a98ed6b939871cb

Observation faa31b6e-e34a-4756-9890-4308bb567c7e · outbound

This paper cites Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.281529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.281529Z digest=sha256:e5753e59e54481ebb334aafd0f495c653dc1e783c3efef47cc1b9249f123bd05

Observation e564af48-2bb2-4c83-a939-f96492fa3324 · outbound

This paper cites Deep q-learning from demonstrations.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep q-learning from demonstrations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.286683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.286683Z digest=sha256:b2db22968debbffda8728da2f8163498b8ade7c8feaa54bde9b970c5a46a2c58

Observation 8ca3dfa6-a5f3-482e-a43a-1373fbd0dd2c · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM 3d-llm: Injecting the 3d world into large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.291172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.291172Z digest=sha256:f0c9bac3b29d46d4a26208a405e566bf9815f69756ef402442bcda785c76f384

Observation 6f6fa83f-5aa2-4d06-8703-9c123f6831af · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.295393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.295393Z digest=sha256:1423bce26807a1b291bd169d17cc1c041dd407599a8a86d09ba835f83bae5f87

Observation d578f68d-f64b-4130-8c0b-34a273fc3554 · outbound

This paper cites Visual language maps for robot navigation.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Visual language maps for robot navigation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.314781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.299703Z digest=sha256:da6569963cea2152231209f46750e50f19e534b987ba9c197192ace4d49b7f8e

Observation adfed43b-9118-4cf9-b9ca-0b365f91c080 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.304138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.304138Z digest=sha256:27c262728fc787484703c6728cf4a399575c963ed0c6cc0ef47758174324f681

Observation 97f35f64-43b2-4862-b800-030416b775f9 · outbound

This paper cites A survey of robot intelligence with large language models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM A survey of robot intelligence with large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.308472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.308472Z digest=sha256:67ec4dc2aaed09999c6a77559d7e4903ca048b6337c418df04a3bfa3a53f7a6d

Observation 9ade9eef-bd05-4d06-82a1-7e3fd16b17a6 · outbound

This paper cites Bench llm deciders with gym translators.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Bench llm deciders with gym translators

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.300138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.313022Z digest=sha256:6cc2411ac19a0f26ed8b554e28799279b58428423a4bead9a419edf8ae19ae7a

Observation aa5fcee6-4bf2-4b9f-a749-6241563b3065 · outbound

This paper cites Housekeep: Tidying virtual households using commonsense reasoning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Housekeep: Tidying virtual households using commonsense reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.285482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.317283Z digest=sha256:837fb60e31fc3f88134e3ed2b2307b51e9db4d905ee740926dde59746beee14d

Observation a3311083-fba8-45a7-b497-08972e02ca09 · outbound

This paper cites Reinforcement learning in robotics: A survey.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Reinforcement learning in robotics: A survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.321708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.321708Z digest=sha256:d8a726ce22eaed30d72fbdf789adff50535b71de3aa2f8917f58d9f79cddf19d

Observation 0fa4467a-407f-4a72-b660-81f06ff2e7fc · outbound

This paper cites Can large language models explore in-context?.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Can large language models explore in-context?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.325869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.325869Z digest=sha256:48c93ada6de31c8cbe4e2f610fa8ac20ca2a05b804c328c7185fc548b722324a

Observation d236be72-1f30-400a-b6f9-fc799da41d2f · outbound

This paper cites Batch reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Batch reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.330290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.330290Z digest=sha256:0fd3c43200248f8dce3ea5cb234d9c05993a64d2b4ef9473f65e5736134f5b36

Observation 41565e31-b181-4d8e-9d97-25c573350a16 · outbound

This paper cites Supervised pretraining can learn in-context reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Supervised pretraining can learn in-context reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.335032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.335032Z digest=sha256:92f186309f9c04ab3f2e73d9f7b680e25aa27c875d85bba297ffd2f12d51a792

Observation 6bc88d04-51be-41ef-98bd-feb84b348992 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.339425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.339425Z digest=sha256:92cc102f30be4fb633251e34f8f5337a74fe485630afac3614bbfdbe49e971ca

Observation 1ac0219f-c239-4f1d-b07b-1b8156ccb9f4 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Code as policies: Language model programs for embodied control

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.344099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.344099Z digest=sha256:114608e5e0527b7f5eae34cfed004bb44edd561168e87fa18a91b87d7ffa15bd

Observation a90147ac-36aa-4152-be0b-12e7a0655000 · outbound

This paper cites Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.348912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.348912Z digest=sha256:5ab00fa391cf2099c29838e75a239dafadd4e6ea3d2bb9dbcdb5b8ae583c5393

Observation 07649708-e6e6-45e2-82aa-4f5d76599e6d · outbound

This paper cites Physgen: Rigid-body physics-grounded image-to-video generation.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Physgen: Rigid-body physics-grounded image-to-video generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.236332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.353765Z digest=sha256:bda5939c1b6cf19aab35e96490559f85cca58b015182a2493c376fdd0da14708

Observation 576cefc7-e47d-47c8-aff9-632f26c4d459 · outbound

This paper cites A Survey of Reinforcement Learning Informed by Natural Language.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM A Survey of Reinforcement Learning Informed by Natural Language

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.358359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.358359Z digest=sha256:379ee1a80041b94564dd29d16c4e9548adfa838d7294f03aa0cfcd607a54a776

Observation be556654-9cbb-4619-8a4f-accd09f8c760 · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.363027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.363027Z digest=sha256:46a557b49dfaf4b21855c79e438ff168023823921bc5316274fa829ae054edf8

Observation 2db7ad14-22a4-48dc-bccd-84606c27c9b8 · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Overcoming exploration in reinforcement learning with demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.367902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.367902Z digest=sha256:a634dbfc7b4b81e872023e897b7847497f7adf833f48c44045d33b314bb0ef50

Observation 52fc9b45-c769-42c5-8c02-f35e96739d8b · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.372358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.372358Z digest=sha256:a671dd6278bd563443e0bdf9e89e976e3e4326ff2483421cc88aec264d6c4059

Observation e94f63ce-d0b5-400a-a8e0-1f0595fdf8e8 · outbound

This paper cites EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.377077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.377077Z digest=sha256:8942b161b396d101d115e229441cc21dd9e22d76e61d8e0832ee08d346285ce0

Observation d095b5cb-644b-4028-8449-86f33522ddb5 · outbound

This paper cites Llamagym: Fine-tune llm agents with online reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Llamagym: Fine-tune llm agents with online reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.213175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.381607Z digest=sha256:deb335da52907a43ffab8367d879e1cdcda4d6ca4271917f2520669b2392c1f2

Observation 6db3aedd-05f8-478f-af60-1ef3638e7622 · outbound

This paper cites Mapping language models to grounded conceptual spaces.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Mapping language models to grounded conceptual spaces

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.385866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.385866Z digest=sha256:cb3086dc9e2cdbc9b12ae545a21d9d9cad376621cb13e9cedb526fcf1c88306c

Observation e3d43e25-5807-4d58-a31f-eca6ac07c11c · outbound

This paper cites Deep Q-Learning: Theoretical Insights from an Asymptotic Analysis.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep Q-Learning: Theoretical Insights from an Asymptotic Analysis

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:05:21.695486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.390484Z digest=sha256:d440463f255dc820120c9d1b2fe64d2799122ee9369bec3e6f79627cf97ec025

Observation 92b22de6-bf68-4b0e-bc43-45cd26ac1e31 · outbound

This paper cites Neural fitted q iteration--first experiences with a data efficient neural reinforcement learning method.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Neural fitted q iteration--first experiences with a data efficient neural reinforcement learning method

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.187685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.395272Z digest=sha256:096787b3bea4f1aa3e74fdef2e8b05d6d8f63c4cca618e6d8f18a40a1ae3f2f5

Observation f695cd43-3bd0-4c65-80d6-19a3efb495fb · outbound

This paper cites Kickstarting Deep Reinforcement Learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Kickstarting Deep Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.399556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.399556Z digest=sha256:9bcb221f57d6fd64144a970605b48c6fc68c5019ef930e2999a489b535ad446f

Observation dd8f0dfc-628e-46a6-a737-0f46f2607295 · outbound

This paper cites d3rlpy: An offline deep reinforcement learning library.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM d3rlpy: An offline deep reinforcement learning library

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.404084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.404084Z digest=sha256:2cf7424f693155b3a8e049dda643fd443f7205f35627c88efe17830f30997b6c

Observation b0e751dd-a5ec-47a2-ac7f-a1c8ee1cecf4 · outbound

This paper cites Mastering the game of go without human knowledge.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Mastering the game of go without human knowledge

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.408938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.408938Z digest=sha256:b1d8ee5233d6cd29ab07fc13559f5d0deb1ffb4ce4bc47b7d3b493857b13aca5

Observation 97cc27ab-56ac-4efe-8d15-7e70b2245994 · outbound

This paper cites It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.413563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.413563Z digest=sha256:4f27d50776196d85666783ffd343b6751b607ab421dc0c4421e74371a398f61c

Observation 143dc137-5a88-49bb-ae91-a05c169d57b4 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.418269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.418269Z digest=sha256:ed828adbcc9216f05432c34b1a2501fee6d8c1db9f800fbd8473e786d457edeb

Observation 8a37a8d8-fd26-42ef-aa7e-625726d1c3a1 · outbound

This paper cites True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.423074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.423074Z digest=sha256:d0a7e65527025cf6825592c66560a2ec31f71413f456995caed1c8481bdfa0f2

Observation 3e024e96-912d-454a-80e5-b4b9abfacfc5 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.427700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.427700Z digest=sha256:aa9b3db89590a068b3198fcaa64dfc44d3164d1b9765d40edc23d09ad2338e8f

Observation 6dcd2c35-d5d6-4a16-ad07-9515d53786e3 · outbound

This paper cites Representation Learning for Online and Offline RL in Low-rank MDPs.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Representation Learning for Online and Offline RL in Low-rank MDPs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.432391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.432391Z digest=sha256:eca3d42f1208aaa9e50735d70c4a5ddee507481cbb01c7a109fa31aef81add99

Observation a59c1136-1d1b-4ce4-8d72-2b6a99b53dbd · outbound

This paper cites Deep Reinforcement Learning with Double Q-learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Deep Reinforcement Learning with Double Q-learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.437109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.437109Z digest=sha256:677e48ef18f2977cf19e8e89a585fa628476d1b68023e8374982eff269d39b85

Observation 0d28139d-0e1b-4166-b820-945784beb028 · outbound

This paper cites Instabilities of offline rl with pre-trained neural representation.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Instabilities of offline rl with pre-trained neural representation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.153897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.441831Z digest=sha256:908f20501da5d579f3e2150b88e3f605ce7d30ea3d6e62b42337e6a39b4e06ff

Observation 1b2f8424-0178-4a98-a6f5-b7161f2540a6 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Chain-of-thought prompting elicits reasoning in large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.445957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.445957Z digest=sha256:029e9188c6dd1abae08a836e1a363a08e6ec4bc6e67bfbfbe792237fcd7e8a48

Observation c4e53891-9749-4d06-a701-4061e7827b72 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.129447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.450587Z digest=sha256:19d48631551bd3bb158dc6d9a9867df9012acc14e2830dab967c6ddb770b8ef1

Observation a61e155c-1586-4e41-af76-13ed7df7292f · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.454824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.454824Z digest=sha256:4ea417fbc0229e1c9d2cdf39574e9103ea393efd52448550bf9e89210b96f967

Observation bd68ada6-c12b-441e-8fe8-4b81eb54573e · outbound

This paper cites Qwen2.5 Technical Report.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Qwen2.5 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.459606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.459606Z digest=sha256:8267cd3e81612f2f58ebb8ece9ddedf0b579b0ea66b30a73dd7e461f94657efc

Observation 9b2129aa-1500-4fb9-b8a8-401021eb3933 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM ReAct: Synergizing Reasoning and Acting in Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.463838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.463838Z digest=sha256:444cb97d30effac04ab1c5b67dcf3119439a686ef4d94bed3860235d1aa29e6c

Observation 41044195-3ce1-4045-a71f-decbeb0082d3 · outbound

This paper cites Policy finetuning in reinforcement learning via design of experiments using offline data.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Policy finetuning in reinforcement learning via design of experiments using offline data

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.114177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.468622Z digest=sha256:406ee6bccfebf196107e2d63a9ca6c4148d5dea2c211f31645e3e28143c97ed6

Observation a4504fb6-8505-49b8-b05a-66dc89d85d30 · outbound

This paper cites Online decision transformer.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Online decision transformer

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:05:22.097815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T21:05:21.472787Z digest=sha256:5733c5eb790d2596f9bd4b32780c7c5bd7a3034ea872e0c468c38a87af1afc42

Observation 2e7e81b9-3538-4909-9f76-5aa26cd04b72 · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:21.477420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:21.477420Z digest=sha256:4b7d0c419b6296808081fa735721c28944c55da3f428b647c276be65260db223

Pith citing papers

Observation f2d30702-5654-451e-a9d5-872b374fdef9 · inbound

ProDVI: Programmatic Dynamics Priors for Value Network Initialization cites this paper.

ProDVI: Programmatic Dynamics Priors for Value Network Initialization Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:25:38.695370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T19:25:37.918588Z digest=sha256:32a8b6f17704c41ae92c41c2d4a6650790d91eb06ae9648ad19c5b0e8358fed8