Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T20:18:05.074641Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2602.23721.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T20:18:05.074641Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1e7f9b65-7ac1-4643-abde-cebf0f1662b3 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation OpenVLA: An Open-Source Vision-Language-Action Model
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbfa63d4-0c29-4e5d-9237-ae9f1f273746 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Joshi, Ryan Ju lian, Dmitry Kalashnikov, et al
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30bdf333-dc13-4753-94d7-3fd28fcf000e · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Embodiedgpt: Vision -language pre -training via embodied chain of thought
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf22649-03f4-426f-89cb-e4ab87cf6680 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Learning man ipulation skills 17 through robot chain -of-thought with sparse failure guidance
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1745af-5682-4d0b-88f1-967fe4873be7 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Robotwin: Dual -arm robot benchmark with generative digital twins
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7652a522-3584-4c3a-a5af-b0e15b82ebcf · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation ScissorBot: Learning Generalizable Scissor Skill for Paper Cutting via Simulation, Imitation, and Sim2Real
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c90874-ddf0-4f14-b451-50dcbd0a5662 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Open x-embodiment: Robotic learning datasets and rt-x models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909cf3f7-a003-4757-8026-73c9a3ac6d0a · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Octo: An Open-Source Generalist Robot Policy
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ac52c3-b240-4701-b864-ccd4b8d6fff9 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Unleashing large-scale video generative pre-training for visual robot manipulation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc9485fa-6219-4cfe-866e-769d957b8ef4 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Cliport: What and where pathways for robotic manipulation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f064819c-76c3-44ea-82c3-1336c4c75be5 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5921040e-8d6d-4937-8a54-344361682922 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fbf0f40-924e-4d9c-88a6-62798d7b7426 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation PaliGemma: A versatile 3B VLM for transfer
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8f7c67c-294d-43c4-abd9-0d4bf78fa1ef · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Vision-language foundation models as effective robot imitators
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8923ce87-c679-4116-b4d4-04d5b2c5308e · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0913954d-f737-4d18-a8d4-c57a5269c5d8 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6098eb69-1345-472c-a857-7143b4121e82 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87aaae01-72b2-4596-833b-9378e6f387a7 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084fe1bf-f199-42c7-be6f-d739ecef2c2b · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Flower: Democratizing generalist robot policies with efficient vision -language action flow policies
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 606138cf-ce84-4313-9db3-340f266a5d47 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b24b7b-723d-4b18-8249-1e491be764bc · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Rt -trajectory: Robotic task generalization via hindsight trajectory sketches
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8804f380-13be-4d91-b5d0-bca1c2acb3ab · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Pivot -r: Primitive -driven waypoint -aware world model for robotic manipulation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 618d9953-7972-404f-886b-27fd379a01ee · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Any-point Trajectory Modeling for Policy Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395e1886-863e-4193-a17a-3b0831c7a8b4 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation DreamGen: Unlocking Generalization in Robot Learning through Video World Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ced7e81-5fed-4373-b3d5-fc5d535c9873 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8cd7914-b628-4f76-a057-a5d7f4368dd6 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20150314-224b-4869-811b-611245e27024 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321b86d1-5eed-46f6-a098-d73b5d0b554c · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94347f0e-5c55-40da-b219-b9dd290ba06a · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Shapellm: Universal 3d object understandi ng for embodied interaction
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b084a0-3c45-48bc-b394-b6f44d0e959c · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Navid: Video -based vlm plans the next step for vision-and-language navigation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f13b23-9f66-44ec-a973-67f421241755 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0726e0ed-c6ae-45ef-9731-0cd6a1677d2e · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation LLaMA: Open and Efficient Foundation Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb56b77-2533-4c5a-8b65-606833ff27dd · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Openaio3ando4 -minisystem card, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de293f5-89e8-4ded-81ed-12a26a464d49 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation DreamLLM: Synergistic multimodal comprehension and creation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01514cb3-0424-4a99-80d8-0ba8944d8162 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c84d68c7-2e32-440c-81b3-6844818041ff · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation VGGT: Visual Geometry Grounded Transformer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b63741-d1b0-4da9-9223-423895f78d7c · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7fd9e90-8197-4df9-aa5e-c31a549271e5 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Tran, RaduSoricut, Anikait Singh, Jaspia r Singh, Pierre Sermanet, Pannag R
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774e5b71-a8ab-4319-bebe-d637b2c8a2f7 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 203de78f-01e3-4723-a92f-d0d450987076 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150d728d-7652-4a3c-bc24-46f9e1ba84ac · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Hume: Introducing System-2 Thinking in Visual-Language-Action Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6076b4c8-a8d6-499a-97b8-44d67fef01f6 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Diffusion policy: Visuomotor policy learning via action diffusion
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f3888f9-e553-430c-aa38-42c8124f6990 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f5aed65-ba22-42e8-810a-6e8f0cf1c86d · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c77afcd-0230-4624-8cc2-b75bbbbdefe3 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation 3d diffusion policy
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81b74bb-33a3-4b6b-8475-0a65ca3cdb35 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Calvin: A benchmark for language -conditioned policy learnin g for long -horizon robot manipulation tasks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa5d9a2-08b7-48f7-bc86-b4584ae0e0c8 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed39de5d-5c00-446f-8612-24bf79080585 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13cf5b98-598f-4cab-8cfc-324181841f03 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc921bac-ba23-479b-bddc-a337264d2d39 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd9bc63-4a2a-4dd1-a853-da566f77e283 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation LIBERO: benchmarking knowledge transfer for lifelong robot learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca160281-d00d-4ac7-878c-a2a8755c02bc · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Decoupled weight decay regularization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c0341f-9bb3-4bbb-ac8b-a4f62381251d · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation OpenAI blog, 1(8):9, 2019
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be7d474-c644-42f4-bac5-c1cbc7faa9e4 · outbound
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Unresolved cited work
Reference 238
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.