Pith. sign in

Paper Citation Record · LEDGER

Improving Vision-Language-Action Model with Online Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2501.16664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16664 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:10:11.693083Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.548397Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4dcedf2-a314-48ff-a734-3fb524645888 · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:38:11.188770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:11f043b48fc91647f3aa85573c6fb3be025367821fd283920bce4f8a6e27b93e

Observation b0f97b08-27e5-4fcf-a420-409b0c1ef202 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:48:48.815167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:8d14315fa5331781e747cec86af41090bbb8b185128def465b2a3619a1851482

Observation 7108ed8c-d69a-4c54-b1fd-20506952726b · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:55:40.499761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:aa3f7d625f8bda6d6dd8008305426133e2477f12f44e24e84c30cbec728405fc

Observation a27cf8f7-cea0-43b4-b7c8-1ec0799ea99a · inbound

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback cites this paper.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:11.693083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:11.693083Z digest=sha256:31844de7b7ea51b561029c14b8a90cce89fc5cbf2dcf5844754acf89f5825b75

Observation 7c594d43-1aa7-4147-8322-d05287c928d0 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:21.447558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:21.447558Z digest=sha256:19bd2790f722d65e5d68859ef05af473a21f2ae105a674422b9b4c713c133e47

Observation 65301ea4-71ca-451a-a691-056ceb3600ae · inbound

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models cites this paper.

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:36:12.864759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:36:12.864759Z digest=sha256:3a2311d2ec2642bd08f813ac25b6d500d834dfdceb156ef055eb78a6a74907e5

Observation 689c3033-cf16-4782-a3a7-eb31f2b0ae21 · inbound

SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training cites this paper.

SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:53:06.290688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:53:06.290688Z digest=sha256:b3b86ce0a974dbd79e5c73e664e88cf09e077bf28256aca5479ffadc587db226

Observation 2634ff7d-e50f-4d87-b330-b96d3689fa0c · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:43.638797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:43.638797Z digest=sha256:aafba24232479a9a3c4e46a3770a63055d564f7bb31ce3f33bade8ffb5af24e6

Observation cd37d08d-cc3e-430e-9226-c59979728f02 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:40.022379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:40.022379Z digest=sha256:c4f8e79995ec2885d3bef8e3bdd1119ad501d88e3371c81bf552cf52e422de4c

Observation 35ba3196-ca77-4c1e-8e83-1549cb75e2ac · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.538701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.538701Z digest=sha256:344d679afb9aa320f7ee601b9c87468a858bae2bb9cfdcf59a01270ddb8a9c7e

Observation 2725aca8-5c7c-48f3-8fbf-57cd903c3e8f · inbound

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization cites this paper.

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:11:31.875241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:11:31.875241Z digest=sha256:13e4428ee17643bd69d5e7c67f4adff58f09a1c15133e17b374f066d8f778d57

Observation 0bfe0573-d8ee-4686-9445-2942dda656c4 · inbound

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning cites this paper.

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:02:11.482284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:02:11.189795Z digest=sha256:a64b820745989c1adfec0c52a9b1f0c9858464df53bf212491e034b8af627bed

Observation 1ab1b8f3-bc2c-4628-bbb2-d0007fb19793 · inbound

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models cites this paper.

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:24:56.027069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:24:56.027069Z digest=sha256:228cf14d418ac0bbe8f51b7c7b1e6b3e5c5e17aea52918c405bf22462f7f4fbf

Observation 98b2c27e-2eaa-4d5d-bc6d-f8c540b37840 · inbound

Ctrl-World: A Controllable Generative World Model for Robot Manipulation cites this paper.

Ctrl-World: A Controllable Generative World Model for Robot Manipulation Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T01:14:10.368029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T01:14:10.174044Z digest=sha256:c99b58ab2581334246c0e6f75a81d717ffc1d0c217d6c2f338561e56f783499f

Observation 90b52234-e401-4d30-be29-95df048fed96 · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:31:02.913540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:ecc5d508bbabfa72e16606ef09c57e50e55080d13c7e842c5be9339a3981248f

Observation 96f1fb71-e18f-481e-972d-a1fafab88cf6 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.347237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:ff10d4d5088838637f765d01cc0424ea51f6805ce5395a64e821a0d773506704

Observation 1ab87388-22a1-49df-b0e9-e41fa555e3f4 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:24.794126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:24.794126Z digest=sha256:5a8f1f7decbf224e7e79aad3096f068e97667881f0b44336acd7c007e2e295ff

Observation 50f1e294-9b47-479e-a108-da46695ada15 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.790260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.790260Z digest=sha256:faf1d4392eea97ffa13828f4b70f687a274314a6375ac927ebefb793c380696a

Observation 58ea4866-2491-401c-8be1-2517f1eb356e · inbound

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation cites this paper.

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:00:25.994640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T07:00:01.741166Z digest=sha256:ad5ef08c77edbabf14a63a083e8becb7ac9dfc669290da64904eb69db8c97839

Observation a8596dfc-c014-4905-aa7b-adc13a2adb55 · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.897478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:6981b7a39fef4b63b6c45894c6575b0ad52b99952c0f64e7b05ca6fd32476b75

Observation 0d17fbb1-bf9e-414c-b44b-a41290644b48 · inbound

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning cites this paper.

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.288701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:07:10.387869Z digest=sha256:b23ef6e3602a141deff7d7161d8f644661ea106fabefa5773e58fb30dd7d9cd1

Observation 694260c5-fe28-4f0e-baf9-87fcd39c434c · inbound

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal cites this paper.

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:26.867422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T01:22:42.691009Z digest=sha256:1008be24948b8be4e8396f007671c91ab50d6dbc6a69d8ec9f9ae71220a4878c

Observation bbc803a6-8cfe-467b-9d40-981a1f8da0fa · inbound

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? cites this paper.

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:10:49.569291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:10:54.362107Z digest=sha256:82a4e18d9a369862642685dcaf3e87611c2fab65f91479d09e009ec3418a34f3

Observation f1ace14f-8eea-428c-b697-830e5a049e2a · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.347579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:47d9f6dc67855c4dc7071d83ee9972235e86d2eacf0f7784c6a19fe223358da4

Observation 44d8d85a-3225-42b0-bb30-1eb410e04121 · inbound

Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking cites this paper.

Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:09:02.659480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T21:05:45.024226Z digest=sha256:9582d657d42626215f15760ff18f70575757c2ae067a4e059764523586193021

Observation 56c18ea0-13ad-4965-98ff-870fb0e1512d · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.257279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:31756ee80c4048985c6edaa56536ed5a994b23f6dd02045147d71988f979e312

Observation 510528f5-e4b1-4196-8011-37b2e07639e6 · inbound

Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models cites this paper.

Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:23.617278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:49:15.137844Z digest=sha256:61386ea58371f92b0085703b3a3ce6e7591227c3f8f7159c8924d374a0256fa5

Observation bbde08bf-32e7-433b-9ce7-afe9dc3ef770 · inbound

AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning cites this paper.

AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.001497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:25:59.194721Z digest=sha256:faf95ff646b70c16e6f087b67ed615bbbb2420f08bd95c80ffd5085d638a18a6

Observation 69d21ad9-a07a-4301-af0e-5d08773fa56a · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.550081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:10ec21573d0d89e9bd1fa8a67a6b1ed347db4c312bbb96166554859d6ba37331

Observation 5a0155eb-bdf8-4494-bf6d-5b47aabf7ef9 · inbound

Adapting Generalist Robot Policies with Semantic Reinforcement Learning cites this paper.

Adapting Generalist Robot Policies with Semantic Reinforcement Learning Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.654679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:09:29.625066Z digest=sha256:0ba1e6ad4bdcb1d6e8456e5d8ad1e6c21e5f23306b6d63921a913f7e04f9dba8

Observation 844fe8d3-167e-4ef6-a8df-f31242600c73 · inbound

RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment cites this paper.

RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:40.214321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:40.214321Z digest=sha256:48ea809d8850c776eee41fe1d6690effa38ed5d62a0f46f059ee8a08605e98fc