Pith. sign in

Paper Citation Record · LEDGER

Improving Vision-Language-Action Model with Online Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2501.16664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16664 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:10:11.693083Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.548397Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4dcedf2-a314-48ff-a734-3fb524645888 · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:38:11.188770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:ae31aed4bcbbd74b85e84df1d3a912df3be926a555adb1878145750337588c0e

Observation b0f97b08-27e5-4fcf-a420-409b0c1ef202 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:48:48.815167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:b2e29f43250c03f54c34f3b2f0823ee387d954b29ef721845e49a6e92e7dc252

Observation 7108ed8c-d69a-4c54-b1fd-20506952726b · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:55:40.499761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:78c12e669c577b992a8ddb281d3ba5d2d38cfb34b86ca10476d1618a5fce10dc

Observation a27cf8f7-cea0-43b4-b7c8-1ec0799ea99a · inbound

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback cites this paper.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.693083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.693083Z digest=sha256:3b7ecfc80a7cdc468390d9800554e77e49c1c022a249563e1ef716347415648d

Observation 7c594d43-1aa7-4147-8322-d05287c928d0 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:21.447558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:21.447558Z digest=sha256:19bd2790f722d65e5d68859ef05af473a21f2ae105a674422b9b4c713c133e47

Observation 65301ea4-71ca-451a-a691-056ceb3600ae · inbound

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models cites this paper.

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:36:12.864759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:36:12.864759Z digest=sha256:3a2311d2ec2642bd08f813ac25b6d500d834dfdceb156ef055eb78a6a74907e5

Observation 689c3033-cf16-4782-a3a7-eb31f2b0ae21 · inbound

SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training cites this paper.

SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:53:06.290688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:53:06.290688Z digest=sha256:b3b86ce0a974dbd79e5c73e664e88cf09e077bf28256aca5479ffadc587db226

Observation 2634ff7d-e50f-4d87-b330-b96d3689fa0c · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:43.638797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:43.638797Z digest=sha256:8566cefe3adf15368b94dc2c7e60ea0ace79bcb4343d0e9e4218fcb56b43c441

Observation cd37d08d-cc3e-430e-9226-c59979728f02 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:40.022379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:40.022379Z digest=sha256:c4f8e79995ec2885d3bef8e3bdd1119ad501d88e3371c81bf552cf52e422de4c

Observation 35ba3196-ca77-4c1e-8e83-1549cb75e2ac · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.538701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.538701Z digest=sha256:344d679afb9aa320f7ee601b9c87468a858bae2bb9cfdcf59a01270ddb8a9c7e

Observation 2725aca8-5c7c-48f3-8fbf-57cd903c3e8f · inbound

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization cites this paper.

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:11:31.875241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:11:31.875241Z digest=sha256:13e4428ee17643bd69d5e7c67f4adff58f09a1c15133e17b374f066d8f778d57

Observation 0bfe0573-d8ee-4686-9445-2942dda656c4 · inbound

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning cites this paper.

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:02:11.482284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T08:02:11.189795Z digest=sha256:b4be7e8b08dd1d21261ec09110248ac23a82df501ed0e7921d9b8dacdb515507

Observation 1ab1b8f3-bc2c-4628-bbb2-d0007fb19793 · inbound

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models cites this paper.

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:24:56.027069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:24:56.027069Z digest=sha256:228cf14d418ac0bbe8f51b7c7b1e6b3e5c5e17aea52918c405bf22462f7f4fbf

Observation 98b2c27e-2eaa-4d5d-bc6d-f8c540b37840 · inbound

Ctrl-World: A Controllable Generative World Model for Robot Manipulation cites this paper.

Ctrl-World: A Controllable Generative World Model for Robot Manipulation Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T01:14:10.368029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T01:14:10.174044Z digest=sha256:007be608747e02836cc1da796240f0f49e55d8646942a3186507a8d26c6281c4

Observation 90b52234-e401-4d30-be29-95df048fed96 · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:31:02.913540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:d9d7a7f0e8b50a6117dd860ebf6525d72d3672fa91b72872d9e64562f1b056d0

Observation 96f1fb71-e18f-481e-972d-a1fafab88cf6 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.347237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:137e2d8596996f7ed22f8ce0c2d4ba37a063523c64c01c077b37ab3898f3ff4e

Observation 1ab87388-22a1-49df-b0e9-e41fa555e3f4 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:24.794126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:24.794126Z digest=sha256:5a8f1f7decbf224e7e79aad3096f068e97667881f0b44336acd7c007e2e295ff

Observation 50f1e294-9b47-479e-a108-da46695ada15 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.790260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.790260Z digest=sha256:faf1d4392eea97ffa13828f4b70f687a274314a6375ac927ebefb793c380696a

Observation 58ea4866-2491-401c-8be1-2517f1eb356e · inbound

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation cites this paper.

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:00:25.994640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T07:00:01.741166Z digest=sha256:dcb75edb3d10999c76222f6972708da71eb847acc8eb5b00b57131b50b981e8c

Observation a8596dfc-c014-4905-aa7b-adc13a2adb55 · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.897478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:b5e2742d23688214cb5f05b74ec429e5b8507db627757c0c75101d7d55f15366

Observation 0d17fbb1-bf9e-414c-b44b-a41290644b48 · inbound

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning cites this paper.

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.288701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:07:10.387869Z digest=sha256:59cc28b868653ee0f103b12551c0636ba9c7f1420533ad9f353f410c0419a0f2

Observation 694260c5-fe28-4f0e-baf9-87fcd39c434c · inbound

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal cites this paper.

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:26.867422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T01:22:42.691009Z digest=sha256:cf6bddecbaa7afafa8ae4267533675854f4107c11521038ba565175a978386ac

Observation bbc803a6-8cfe-467b-9d40-981a1f8da0fa · inbound

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? cites this paper.

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:10:49.569291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:10:54.362107Z digest=sha256:8e4721c26691c08069ac35a8adc4335d46464943b63316eaa922e9925cc493a7

Observation f1ace14f-8eea-428c-b697-830e5a049e2a · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.347579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:81629655a58d2ff681265653b0e93e38969f9da74f50d7551a3008a2a95afdb8

Observation 44d8d85a-3225-42b0-bb30-1eb410e04121 · inbound

Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking cites this paper.

Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:09:02.659480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T21:05:45.024226Z digest=sha256:0bd4ee21b931bc0b12e5bcc263a0f3522329d745f6052c0ef72754d2c4367a21

Observation 56c18ea0-13ad-4965-98ff-870fb0e1512d · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.257279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:9629ec0c68da83fea94bb0ff54ebff6f4613e5a2fd1bd82628e8b426c5ceb112

Observation 510528f5-e4b1-4196-8011-37b2e07639e6 · inbound

Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models cites this paper.

Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:23.617278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:49:15.137844Z digest=sha256:b86d54b8b86ff796e945ce0d841539dec20f1e8ddaf6da6af43de82ad080f960

Observation bbde08bf-32e7-433b-9ce7-afe9dc3ef770 · inbound

AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning cites this paper.

AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.001497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:25:59.194721Z digest=sha256:f8dadea7b5ae1971e85118742ff2cf9c74b4595dc97a05e91477d38e6f81dd43

Observation 69d21ad9-a07a-4301-af0e-5d08773fa56a · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.550081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:9f253321dd928573d18804295f80894bad6ded558ca31db1ff0b3607ebdb09f9

Observation 5a0155eb-bdf8-4494-bf6d-5b47aabf7ef9 · inbound

Adapting Generalist Robot Policies with Semantic Reinforcement Learning cites this paper.

Adapting Generalist Robot Policies with Semantic Reinforcement Learning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.654679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:09:29.625066Z digest=sha256:b948af06a97f49a7d15f66ddddb49909d5f7574af3f8a884a53d298d580dc3d8

Observation 844fe8d3-167e-4ef6-a8df-f31242600c73 · inbound

RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment cites this paper.

RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:40.214321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:40.214321Z digest=sha256:48ea809d8850c776eee41fe1d6690effa38ed5d62a0f46f059ee8a08605e98fc