Pith. sign in

Paper Citation Record · LEDGER

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

As of 15 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 35 inbound Pith citation observations for arXiv:2501.18867.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18867 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:13:14.953745Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:28:25.098888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:19:38.145306Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77eed692-b15f-48b4-a948-a61c91b532c8 · outbound

This paper cites H., and Krishnan, R.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent H., and Krishnan, R

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.865706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.865706Z digest=sha256:1adeb04c691dac59881eceda8a3db6004b5eb334370c4242f568922e4215b216

Observation 8dbec790-398b-4e18-97f4-da94819ecf35 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.881237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.881237Z digest=sha256:3c26dce869802a1e5b20860f6fdef5b30abb12b55d03619842797d074e868f82

Observation 3d99890c-6af2-4d2e-94e9-b759edc8c961 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.870144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.870144Z digest=sha256:766ec7498953d14139a99be1a7d097c63805856baa9e7584a8b31a10e6252ea2

Observation f78bd75c-6e8c-4692-8777-791cbb8c8c34 · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.887809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.887809Z digest=sha256:8dd7a0ae9e7cc7e36dd86a7afc130588563a863cca557b8f012e3d0bd2d4c125

Observation 6dad2403-9e6c-4961-8cd5-023ef2f1750e · outbound

This paper cites OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.890752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.890752Z digest=sha256:1415fdb12ae99176138a6b6f41ad4834672f071acdf0064ccec7ef8a38ca3d4f

Observation 0d62aff4-d491-4305-9aa0-840ef2e1fef0 · outbound

This paper cites Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.893941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.893941Z digest=sha256:2cb8ef7244dbb7e6b5dbccb383af69d1a6f25d6aacc95447b9b6494deb3fbe89

Observation 90af2f4f-1696-482c-84af-1a7fd36358b2 · outbound

This paper cites Prediction with Action: Visual Policy Learning via Joint Denoising Process.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Prediction with Action: Visual Policy Learning via Joint Denoising Process

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.907035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.907035Z digest=sha256:5687003c2e46f6b7158b5238c0162ff5a4cb3f41e545c0f9f65f57930d5944a8

Observation 553fc887-2af7-46bc-90c0-23c0867ebc63 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.910522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.910522Z digest=sha256:66de9ef69c60cc6d9a6ad462519621d592878a5ce7ebfdd6cf7d50693a5c1806

Observation 699109ef-510d-4910-b3df-9cb8b1a9da89 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent OpenVLA: An Open-Source Vision-Language-Action Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.913213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.913213Z digest=sha256:f57e4c5a4829b133deae3311a106dbe705334e1effc24468091a9e34cb7a96c3

Observation 55ed24c6-9829-42df-af65-7ad72d392227 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Vision-Language Foundation Models as Effective Robot Imitators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.916949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.916949Z digest=sha256:f6a51eb1d9397527a764ef86073717c1a45a69a759c35338a172eed5aee3f73e

Observation f3f94546-5816-4fac-b698-69929c7201dd · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.919949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.919949Z digest=sha256:cf69291ec7b0e846e01b8b9f9361c8e31433bdc16d83d289c60a9d3cbfd669b8

Observation 623b1317-64fc-4abc-8b21-ebd5b20cb4ec · outbound

This paper cites Accelerating vision-language- action model integrated with action chunking via parallel decoding.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Accelerating vision-language- action model integrated with action chunking via parallel decoding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.922693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.922693Z digest=sha256:69eb2106ecdee5c242f45c3231980bebd054628b4590bcd31a704b0ae1e50707

Observation fe9d8718-a2b9-42bd-a1d4-48c7ef5940a9 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.925841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.925841Z digest=sha256:ec81b55f2174ad40666954a5d7f0ade05ac37c7d623e798ba22a9b4b593a6760

Observation 2d464b53-e88e-47cb-9843-f6611725bd91 · outbound

This paper cites Prompt a Robot to Walk with Large Language Models.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Prompt a Robot to Walk with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.928657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.928657Z digest=sha256:09b3f5a938653ecdeed4adc9fcb0023354047e68cd1c04dcb4e48571e9937b15

Observation a3b1ce49-df06-450e-a414-5f1214e1d452 · outbound

This paper cites Can Transformers Capture Spatial Relations between Objects?.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Can Transformers Capture Spatial Relations between Objects?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.931442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.931442Z digest=sha256:6a2633d7b45e6e8b11c92d6ce411910f4f11cdd86c84345cfa3f1cf13bd38806

Observation e083b778-7dab-482e-9eaa-034822cd5605 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.934241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.934241Z digest=sha256:21aba2265d0158880996de2f6513484e14cbb4e6df3e61cdaa233d1dc4a272e6

Observation a9781eb8-61ae-479b-99b4-4a922375f0e5 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.936924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.936924Z digest=sha256:2675e5d99f2f6a8f7467378bb2ce013bd33c703f991bc10829a88532430a8efd

Observation 99fb6199-f6c6-47af-84d5-8ca35ad1d99f · outbound

This paper cites HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.939639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.939639Z digest=sha256:84aa7d6d8d09a4bfc093485dd3ac22371c89536e5df4ff67094f16b9b4c0c0fa

Observation 9e98d851-58a3-499e-ae3a-cf331b776022 · outbound

This paper cites VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.942425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.942425Z digest=sha256:22d7e81a50810f26d83c9d4b81cdc7825819d39583f2fb47ea00c92135c1cc57

Observation 73dd6cfd-4145-4c5a-acc3-baefae20e169 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.945152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.945152Z digest=sha256:62ec0309742a361b2787f93bbd682ff7cee09546aec97e999ad370e1b9b3f765

Observation 02e68723-079a-4f88-9eea-7a6834eb22ec · outbound

This paper cites LM4LV: A Frozen Large Language Model for Low-level Vision Tasks.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent LM4LV: A Frozen Large Language Model for Low-level Vision Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.948139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.948139Z digest=sha256:3560af7620afa0bc7fe2a99d742997087fe7815795cee92ead8aec18186abffb

Observation 4e8b2b7d-6acd-473d-b090-163027680419 · outbound

This paper cites an unresolved cited work.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T22:13:15.458913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:13:14.951056Z digest=sha256:6b764744c150ae83268fe427366a9a2dc805136d0ad11fa30f0a3a5161498f04

Observation 04a94bc3-df9d-4118-b767-c6e296a3f6a4 · outbound

This paper cites In the pretrain stage, we train UP-VLA for 20k steps with batch size of 64 on future prediction and vision-language understanding tasks.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent In the pretrain stage, we train UP-VLA for 20k steps with batch size of 64 on future prediction and vision-language understanding tasks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:13:15.450405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:13:14.953745Z digest=sha256:3669fdcfd6c0838660fe7594b37c3975fe6c5e55416bfb5f54d88e5608a262ed

Observation 7b03bc9e-3174-43f4-900d-afce1b94102a · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent PaLM-E: An Embodied Multimodal Language Model

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.900822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.900822Z digest=sha256:2ebb755541fc260fc4a7ba2e1db824952f48525f93ccc8bdf3ee7e3e59d12f65

Observation 6ed69bac-7427-4159-8b12-0d9b8f1089d6 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.904105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.904105Z digest=sha256:913977e05801971a7ee6a73acc16b2e487693cf57d14d8bfe1102d1e89d07aa4

Observation 5fe9fc33-9155-48f2-9798-ee2289527e9b · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.884489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.884489Z digest=sha256:ca49de3801894d0dc6315af5a801516e37654d77a2d9ccdde570f2b91c397d1a

Observation aab743bb-bcb1-406d-9ee6-346babc795ae · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.873806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.873806Z digest=sha256:a1f2b0ce27782c396f3f4be8b5881a5937586e63e1e41ec91fe300abd4204007

Observation 2e2ae1f0-01e1-4d1f-9c33-b3b3e0b42e19 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.877303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.877303Z digest=sha256:452e2922bf65beef4f3b3f955554eb52b0013a1ae37549a323985b40165c6f83

Observation a06a7741-d030-48f9-9def-c6d402693429 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.896801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.896801Z digest=sha256:1f50c5142781bb2245547d7d367ec3d07d60a05f95da06eb31c4cd829f72e953

Pith citing papers

Observation ecc36a3e-f8ec-4747-9324-290292a2c573 · inbound

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation cites this paper.

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T10:56:57.921843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:56:57.921843Z digest=sha256:a2c4ba0583e85908d3c4a0d512f32dc1b96466f1cdd72c814c60f1d47c3292f1

Observation 19898cd3-2219-4907-9f3d-17d2f2cee578 · inbound

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning cites this paper.

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:17.797242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:17.797242Z digest=sha256:a0720dbae35d4bca6b5608533122ff861fbcab0d9930b3256ed1483ffad53d18

Observation 7533ed72-c049-4a3a-aad1-b597830c97b6 · inbound

Unified Vision-Language-Action Model cites this paper.

Unified Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T18:28:25.098888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:28:25.098888Z digest=sha256:2cf3f641bb1fc8a4870982be3546395ff62a7c8210b4871cfa6f93dec41a4da4

Observation 0bbf6716-25e6-401d-bc0f-69f035bbb1ee · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.525266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:3d5cc32149569be056a5274eb7237f7464a525871d9ac5a4f047ef36d4f70f36

Observation 1834e392-4731-4b51-b42b-5e7fecf03bef · inbound

Improving Generalization of Language-Conditioned Robot Manipulation cites this paper.

Improving Generalization of Language-Conditioned Robot Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:00:00.281156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:00:00.281156Z digest=sha256:676a0f7f5e78d8834be5b355072b2069d6b5b7248b021ddab5af25449a28190f

Observation 5017bbf7-2e89-4a9d-94ce-046ac81d647c · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:47.885219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:47.885219Z digest=sha256:1daa58f5663e16932680a6ddb5ee6691a1bf860a9d77538896a1b2561146e86d

Observation 13d393e9-8b69-4fe8-8a73-c3a3660eb217 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:45.491893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:45.491893Z digest=sha256:98374c7a3c47f8df92573fcd394239023b6009c1370b53c6abf8172bc6343b00

Observation 1cea0819-f384-45cc-824d-f912b80e0671 · inbound

Ctrl-World: A Controllable Generative World Model for Robot Manipulation cites this paper.

Ctrl-World: A Controllable Generative World Model for Robot Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T01:14:10.412898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T01:14:10.174044Z digest=sha256:bbe56011c63e34a1ede971d9482f694873a7e08fda57e084903949a0a546c398

Observation 3d7c705f-3896-4d61-92be-52552f46fe2b · inbound

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models cites this paper.

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:01:20.326259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T23:01:13.910539Z digest=sha256:205ddcc4ce14fce076f89bfed3594b509fd158be1edd52ff91efae6d825c4841

Observation 82e66e5a-afd9-4728-95eb-7f7b46bde77d · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.889361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.889361Z digest=sha256:184e6b05b7a4159fbf6c9c39d9ebd14df73140cac421b20834db8b3d437a9e80

Observation fc62d89c-84fb-4888-a78c-81b968be093b · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:02.045723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:ab7d3fbc0354ec1f3ce647eb95e13122cfb377d196e2ca719eeabcb6cc7e8fa2

Observation a5932355-278d-421b-aec1-1ead141326b3 · inbound

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation cites this paper.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.167790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:c4b5020eff41e8543cc85cbf42bb8c244364dec3f18870604877e173c3cea5ed

Observation c8cd7914-b628-4f76-a057-a5d7f4368dd6 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.971064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.971064Z digest=sha256:eb43124ae7bb73dc6b8ef1c8931dfdf3b08f5fb7926f7557b54f4ce24d95c922

Observation 8c502b30-011a-4097-8701-a4f775628619 · inbound

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance cites this paper.

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:28:23.112425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T00:27:35.706107Z digest=sha256:5550b122233ee88dea79a021a26062b05b6724c71b5cb060c885ee72d9229f84

Observation fe21e7e9-76b4-4ea2-bf6f-c929037bf61d · inbound

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching cites this paper.

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:59:34.036536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T22:59:21.473359Z digest=sha256:412d7758038b6c09dca6e943f0d0bf2f37f7d20962fa9c4e699c7c6ff9b1b82d

Observation 215344d3-9b62-4ee6-90c6-e167ac1d2ca6 · inbound

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? cites this paper.

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:10:49.501816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T20:10:54.362107Z digest=sha256:39ab9ca16c5556a78efda726852a82bd9e829f45145dd2fbf6254be0cf78e3cc

Observation 6a9b6c95-3c09-4cdd-bcbb-b2b4d9df7eac · inbound

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning cites this paper.

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:48.368301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T05:06:38.517652Z digest=sha256:3c1b12fb0efef66d9aebf2c0c5462e42a34675adde59f975fed3fe6b3644739c

Observation bbe15e70-d9c6-4266-b018-17bf1e4020f0 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:53.701218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T03:39:41.090350Z digest=sha256:c9e5e70a50a0ffd577f2166247b51130ec38d2bfaad95bdd8395846d7d4b376a

Observation a546ed1d-8b0c-41fb-96f5-371f64fb4702 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:26:29.199579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T02:53:54.608425Z digest=sha256:153e02339c8d4d60c955a0b18276196b1f8adee7323fc843a87bf254b78487b2

Observation 1e489c8a-6af9-4c50-bf3b-8b25fb7bb44d · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:19:49.907080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T06:16:48.180290Z digest=sha256:987e1adad8d3342af1356836df75d9d70d4358fc865b2b5d6ebd381489ebfb25

Observation eed7992c-03d2-4756-9207-b9e06610f666 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:26.881607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T04:14:54.885244Z digest=sha256:8bc096c402bd54e9cf2d8a10397f484ccac3dba55c093ef1be7da72c359fb9dc

Observation 06681ab4-54d2-4ec5-9873-e7dee94d86a3 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:17:59.541966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T21:14:56.501485Z digest=sha256:da62e3649644c595a5da2936d848349aecf65461541e18777eb91ca09ba3247f

Observation 8a60aa44-b1bc-419a-abb0-009de18a24d1 · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:25.755642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T04:08:43.222818Z digest=sha256:3bc0196f09a751b0ae431a18fa2d628b9a8e1c01695b2783c379ccb1321b7c24

Observation bbf4990f-9cc8-4f36-8c36-7d5d8f5c96cf · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:25:59.846732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:25:59.846732Z digest=sha256:b0e0a73cb803cd1c531bfbf23c8c3b80019295d5fcc83e20d2d37a19526650e7

Observation 244fea20-afda-4dff-920c-68fe5a1cf106 · inbound

UAM: A Dual-Stream Perspective on Forgetting in VLA Training cites this paper.

UAM: A Dual-Stream Perspective on Forgetting in VLA Training UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.072771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T19:24:57.339949Z digest=sha256:db3022ba9b78bc8993b45fcf3ac328c611ed54da07a00c3ee8dceaee77a13dd4

Observation a7ccadef-7dfe-4e6e-8a80-ae221c8a6dec · inbound

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control cites this paper.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.582695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:ee730cdfff2779c7671eed9c5abfea9c91eca3669708cc7332ae6db7675f4931

Observation a635d8f1-ac24-40cf-a4fe-0205460e94f5 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:06:08.671326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T06:05:52.461348Z digest=sha256:817dca34f522696d04894d3576d10e8854fc574264ff6e995b4ef3c8cd799156

Observation d45834e8-781a-4585-a3e2-d40fe382e7bb · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.304885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T16:52:05.565650Z digest=sha256:3990698ccc446d58fe9905a9cccb7e94356e40c1c29ce4b8fabe430a21d5c40f

Observation 0e610237-103c-4e29-85ab-bfb040883cc3 · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.812492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:59c08017c38d622e3fa23b3ae54b92a9b0d9d73effc4cc71e0fc84ff47d8eb81

Observation 0c011b33-cf82-4e6a-acac-4920df51827d · inbound

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding cites this paper.

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:59.243402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T01:23:02.576098Z digest=sha256:5dbe988337a64a3b14140c3985b78cfb4bad85cd92095222bea9baadec42ee51

Observation b5e73954-1eb4-493a-b5d1-a114cc96fe0c · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.287276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:49b7247f7d1b599dd4afaeb9735c38d4b6bed34ed6d53be47f658c393eda2c56

Observation 5fe65f78-e8ee-401e-8667-7858116c64e6 · inbound

MV-WAM: Manifold-Aware World Action Model with Value Augmentation cites this paper.

MV-WAM: Manifold-Aware World Action Model with Value Augmentation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:38.147129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T14:36:42.051049Z digest=sha256:66d82539486b1c940e05b7433e74c07c7cc118c17b31cf88fab2b11ed390cac0

Observation e6920ae1-14bf-43af-b651-170eb1996765 · inbound

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model cites this paper.

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:21.013976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T07:01:48.330369Z digest=sha256:73ba81f266d95a0673705e7a1f427876e9fa2a052db42efa0c70668e9661a558

Observation 729181ee-8d1a-40f3-ad80-66375c7f5081 · inbound

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference cites this paper.

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-15T14:46:04.341192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:46:04.341192Z digest=sha256:43238b01b7547dd1dfa7eaf39ff6da23db361e2482e8dfc604ecb9839a7deb3b

Observation 76e909f4-f93f-4e90-991a-6a937ec19a6b · inbound

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models cites this paper.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.356467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.356467Z digest=sha256:b98b7121b93ebe891a8bf93e9f49cd907132cc4ff5123a75f219f5dff0663268