Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:11:53.760420Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 12 inbound Pith citation observations for arXiv:2506.00411.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:11:53.760420Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T05:32:20.472823Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T11:08:03.086630Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4a2af5aa-9214-4678-b52c-ff461c051d81 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77ad4cb-596c-4755-9d43-0fc6784262a7 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35: 23716–23736, 2022
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611eeb08-4644-4e15-a778-1ea7082f4471 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff18067-3996-452b-bbe3-4b916fb73641 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks RT-H: Action Hierarchies Using Language
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2625f32f-eb50-44d3-b841-6476248aa767 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks PaliGemma: A versatile 3B VLM for transfer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68be341-bf14-4d2a-a232-7818b06184ca · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95deca43-e5e2-464d-9bda-785345e02078 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e8f96c-6bd4-4f52-ad50-8481b3b626d0 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bce84f4-9a64-4654-ab50-1a81f0449008 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e027e3ac-455f-4719-bcb7-ec0c4b1b4b45 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396dad35-108d-4f25-beea-2953d1be9b1d · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks NaVILA: Legged Robot Vision-Language-Action Model for Navigation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1948774a-2fca-4c78-93aa-6210970583db · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3fc4e88-3361-40d0-9bef-a5335e1f2455 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Palm-e: An embodied multimodal language model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f21d51c6-ae7e-4ddd-b8f0-221e2238bedd · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks A survey of embodied ai: From simulators to research tasks.IEEE Transactions on Emerging Topics in Computational Intelligence, 6(2):230–244, 2022
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 494d03c6-01c8-4a84-bd5f-6d0f74311564 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d22152-43ea-4d7b-a36d-d572447164f7 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c15d737-1812-4af8-ba9e-405a1fe6fea4 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3a4d0f-abc1-438d-8dbb-3e89c20b1864 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f686ef96-63b8-4e2e-bac4-d26f9da7dd92 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Inner Monologue: Embodied Reasoning through Planning with Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 557ec476-913e-47e3-8b7f-4c041f2747f7 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c84b103-0336-4f5d-8135-47005bbcd31d · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2): 3019–3026, 2020
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 820a1bc0-e137-4067-9af1-de82658f6d3c · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks VIMA: General Robot Manipulation with Multimodal Prompts
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65712fe7-bfe9-44b7-b52b-5f09f8b515d4 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Prismatic vlms: Investigating the design space of visually-conditioned language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2a4c9e-8674-4c17-9c5b-62aca763dcf9 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks OpenVLA: An Open-Source Vision-Language-Action Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e511f46-d589-4f42-b24f-5c5beb0ba397 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Interactive Task Planning with Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b256970-321e-4233-a30c-d662d2bc37dd · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae26849b-d5ab-43d1-9fc1-9ea2f5737a9c · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6a8cf29-9d17-40d7-bbfd-907b591e9eae · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Code as policies: Language model programs for embodied control
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cedec1a-53b2-46fd-b1a4-90fde3db3aa3 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b24dac-6fca-4be1-be34-5526efb36c12 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3386d76e-45b7-452b-8eaf-5b90a4ccebaf · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Decoupled Weight Decay Regularization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27dcc32e-f1f9-41ff-82c6-e4dcf848fd98 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Embodied Long Horizon Manipulation with Closed-loop Code Generation and Incremental Few-shot Adaptation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19d7957c-ead7-4c6a-af9c-1d95c5e5b4f9 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926c4d49-089e-447c-ae47-79f582e9e696 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8482fb7b-520f-41f1-9783-1f9b55ce32c6 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Learning transferable visual models from natural language supervision
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf29dd0-4804-45d4-b563-50b5267c01ba · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e9729c2-9d0d-48c4-a8b2-d8135faa509e · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Yell At Your Robot: Improving On-the-Fly from Language Corrections
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a57f0c-fb57-4e52-a7ed-b5d87fcc5baa · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Cliport: What and where pathways for robotic manipulation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ae627e-101d-4083-93a9-756881a0c8a7 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Octo: An Open-Source Generalist Robot Policy
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26fcc7d7-4dd4-4a28-b637-5c9c9505445e · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1333bbd0-2651-4b57-b84e-918508fb06b2 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks KARMA: Augmenting Embodied AI Agents with Long-and-short Term Memory Systems
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 267028b0-7e69-49df-8662-1eedb1064a11 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Chain-of-thought prompting elicits reasoning in large language models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ecbe182-66c7-4eb9-b7b9-233a0db0521f · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d1168fb-ecff-4995-826c-eccec69fc6f6 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0777f90-df48-4f56-927b-3d6696670cd4 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Embodied Task Planning with Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dadc5fd-e2b5-4f32-a48f-cc3b5ba8e226 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Guiding Long-Horizon Task and Motion Planning with Vision Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d053f5ef-cf53-42ef-8fd1-0ad9ef18c56b · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ab15d2-1cd4-40ce-95ac-854b15e2e5fd · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Robotic Control via Embodied Chain-of-Thought Reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32a1e977-58cf-4ec1-aef5-422dc7c48f00 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Transporter networks: Rearranging the visual world for robotic manipulation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029dc6b2-2290-4ed4-a692-683e44f7cde3 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks LoHoRavens: A Long-Horizon Language-Conditioned Benchmark for Robotic Tabletop Manipulation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04cad5a2-d5bb-4cd1-aa82-961911dc6ef4 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0fcf83-bbca-43ca-b251-01076d826155 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Erra: An embodied representation and reasoning architecture for long-horizon language- conditioned manipulation tasks.IEEE Robotics and Automation Letters, 8(6):3230–3237, 2023
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b5230cb-1021-4eb9-8fd9-e7342d94e7e3 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5853545-59f4-46b9-b4bb-6d32d1843908 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks Isr-llm: Iterative self-refined large language model for long-horizon sequential task planning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2fe3bc54-3dad-47d6-97a8-3fb0fa87f083 · outbound
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997d3490-1325-4a65-8c76-7513290a445c · inbound
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8d5e4cc-5464-4f4e-af06-f03fa3161f38 · inbound
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fec2673b-f60e-44eb-bda3-4d0186d6a3ea · inbound
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 828b5d19-bd5e-441f-8931-da55b7502848 · inbound
Open-Loop Planning, Closed-Loop Verification: Speculative Verification for VLA LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6dc473a3-1204-429a-bf56-f4debac64faa · inbound
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc4d610d-b7e9-4f44-91bd-fa748b1f2c2e · inbound
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e491ca-6df3-4e6a-8276-241a1931768e · inbound
ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6720793-1822-43d3-89f0-542ec16861f6 · inbound
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb1f5b37-0eb1-4e8f-8cfb-d3a8159b8376 · inbound
Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 280
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 402cf453-63ce-463a-b2f4-a06344575e19 · inbound
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners? LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e08cc78-0e59-4138-aeab-62b2d17b03af · inbound
$\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6307be61-d675-45f9-a719-c1b73bd6ef7b · inbound
LENS: LLM-guided Environment Simplification for Planning and Control in Clutter LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.