Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:13:14.953745Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 35 inbound Pith citation observations for arXiv:2501.18867.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:13:14.953745Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:28:25.098888Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:19:38.145306Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 77eed692-b15f-48b4-a948-a61c91b532c8 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent H., and Krishnan, R
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dbec790-398b-4e18-97f4-da94819ecf35 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent RT-1: Robotics Transformer for Real-World Control at Scale
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d99890c-6af2-4d2e-94e9-b759edc8c961 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78bd75c-6e8c-4692-8777-791cbb8c8c34 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dad2403-9e6c-4961-8cd5-023ef2f1750e · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d62aff4-d491-4305-9aa0-840ef2e1fef0 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90af2f4f-1696-482c-84af-1a7fd36358b2 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Prediction with Action: Visual Policy Learning via Joint Denoising Process
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 553fc887-2af7-46bc-90c0-23c0867ebc63 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699109ef-510d-4910-b3df-9cb8b1a9da89 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent OpenVLA: An Open-Source Vision-Language-Action Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ed24c6-9829-42df-af65-7ad72d392227 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Vision-Language Foundation Models as Effective Robot Imitators
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3f94546-5816-4fac-b698-69929c7201dd · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623b1317-64fc-4abc-8b21-ebd5b20cb4ec · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Accelerating vision-language- action model integrated with action chunking via parallel decoding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe9d8718-a2b9-42bd-a1d4-48c7ef5940a9 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d464b53-e88e-47cb-9843-f6611725bd91 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Prompt a Robot to Walk with Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3b1ce49-df06-450e-a414-5f1214e1d452 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Can Transformers Capture Spatial Relations between Objects?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e083b778-7dab-482e-9eaa-034822cd5605 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9781eb8-61ae-479b-99b4-4a922375f0e5 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fb6199-f6c6-47af-84d5-8ca35ad1d99f · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e98d851-58a3-499e-ae3a-cf331b776022 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73dd6cfd-4145-4c5a-acc3-baefae20e169 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02e68723-079a-4f88-9eea-7a6834eb22ec · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent LM4LV: A Frozen Large Language Model for Low-level Vision Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8b2b7d-6acd-473d-b090-163027680419 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 04a94bc3-df9d-4118-b767-c6e296a3f6a4 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent In the pretrain stage, we train UP-VLA for 20k steps with batch size of 64 on future prediction and vision-language understanding tasks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7b03bc9e-3174-43f4-900d-afce1b94102a · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent PaLM-E: An Embodied Multimodal Language Model
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ed69bac-7427-4159-8b12-0d9b8f1089d6 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe9fc33-9155-48f2-9798-ee2289527e9b · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab743bb-bcb1-406d-9ee6-346babc795ae · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e2ae1f0-01e1-4d1f-9c33-b3b3e0b42e19 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06a7741-d030-48f9-9def-c6d402693429 · outbound
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecc36a3e-f8ec-4747-9324-290292a2c573 · inbound
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19898cd3-2219-4907-9f3d-17d2f2cee578 · inbound
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7533ed72-c049-4a3a-aad1-b597830c97b6 · inbound
Unified Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bbf6716-25e6-401d-bc0f-69f035bbb1ee · inbound
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1834e392-4731-4b51-b42b-5e7fecf03bef · inbound
Improving Generalization of Language-Conditioned Robot Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5017bbf7-2e89-4a9d-94ce-046ac81d647c · inbound
Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d393e9-8b69-4fe8-8a73-c3a3660eb217 · inbound
Leveraging OS-Level Primitives for Robotic Action Management UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cea0819-f384-45cc-824d-f912b80e0671 · inbound
Ctrl-World: A Controllable Generative World Model for Robot Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3d7c705f-3896-4d61-92be-52552f46fe2b · inbound
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 82e66e5a-afd9-4728-95eb-7f7b46bde77d · inbound
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc62d89c-84fb-4888-a78c-81b968be093b · inbound
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a5932355-278d-421b-aec1-1ead141326b3 · inbound
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c8cd7914-b628-4f76-a057-a5d7f4368dd6 · inbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c502b30-011a-4097-8701-a4f775628619 · inbound
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fe21e7e9-76b4-4ea2-bf6f-c929037bf61d · inbound
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 215344d3-9b62-4ee6-90c6-e167ac1d2ca6 · inbound
Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a9b6c95-3c09-4cdd-bcbb-b2b4d9df7eac · inbound
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bbe15e70-d9c6-4266-b018-17bf1e4020f0 · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a546ed1d-8b0c-41fb-96f5-371f64fb4702 · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e489c8a-6af9-4c50-bf3b-8b25fb7bb44d · inbound
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eed7992c-03d2-4756-9207-b9e06610f666 · inbound
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 06681ab4-54d2-4ec5-9873-e7dee94d86a3 · inbound
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8a60aa44-b1bc-419a-abb0-009de18a24d1 · inbound
UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bbf4990f-9cc8-4f36-8c36-7d5d8f5c96cf · inbound
UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 244fea20-afda-4dff-920c-68fe5a1cf106 · inbound
UAM: A Dual-Stream Perspective on Forgetting in VLA Training UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a7ccadef-7dfe-4e6e-8a80-ae221c8a6dec · inbound
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a635d8f1-ac24-40cf-a4fe-0205460e94f5 · inbound
From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d45834e8-781a-4585-a3e2-d40fe382e7bb · inbound
From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0e610237-103c-4e29-85ab-bfb040883cc3 · inbound
GEM: Generative Supervision Helps Embodied Intelligence UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0c011b33-cf82-4e6a-acac-4920df51827d · inbound
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b5e73954-1eb4-493a-b5d1-a114cc96fe0c · inbound
World Pilot: Steering Vision-Language-Action Models with World-Action Priors UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5fe65f78-e8ee-401e-8667-7858116c64e6 · inbound
MV-WAM: Manifold-Aware World Action Model with Value Augmentation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e6920ae1-14bf-43af-b651-170eb1996765 · inbound
Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 729181ee-8d1a-40f3-ad80-66375c7f5081 · inbound
Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e909f4-f93f-4e90-991a-6a937ec19a6b · inbound
TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.