Pith. sign in

Paper Citation Record · LEDGER

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2602.23721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.23721 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:18:05.074641Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e7f9b65-7ac1-4643-abde-cebf0f1662b3 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:03.450845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:03.450845Z digest=sha256:02b4795f0526cc62384259b7ead5d3362eb5c41268c9218bcee1bd095d997921

Observation cbfa63d4-0c29-4e5d-9237-ae9f1f273746 · outbound

This paper cites Joshi, Ryan Ju lian, Dmitry Kalashnikov, et al.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Joshi, Ryan Ju lian, Dmitry Kalashnikov, et al

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:03.508441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:03.508441Z digest=sha256:eb88cb2a54a47474f473e69f83c9fe1604654e67b0ff765ee6e96850070281fa

Observation 30bdf333-dc13-4753-94d7-3fd28fcf000e · outbound

This paper cites Embodiedgpt: Vision -language pre -training via embodied chain of thought.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Embodiedgpt: Vision -language pre -training via embodied chain of thought

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:03.581910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:03.581910Z digest=sha256:bf11af5f4494fa529fac0368a00d4405529602179d3c6ccb30ea0efe420c8672

Observation 9bf22649-03f4-426f-89cb-e4ab87cf6680 · outbound

This paper cites Learning man ipulation skills 17 through robot chain -of-thought with sparse failure guidance.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Learning man ipulation skills 17 through robot chain -of-thought with sparse failure guidance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:03.610428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:03.610428Z digest=sha256:c78dcca0e032767259491867f2a41405c3bbb1485d622b92531b73f91b5322db

Observation 2b1745af-5682-4d0b-88f1-967fe4873be7 · outbound

This paper cites Robotwin: Dual -arm robot benchmark with generative digital twins.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Robotwin: Dual -arm robot benchmark with generative digital twins

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:03.676641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:03.676641Z digest=sha256:3e6ab72b2a1335437a50b94e4fad60d5b21c34ea37ee2d6a02ca29320df7bf97

Observation 7652a522-3584-4c3a-a5af-b0e15b82ebcf · outbound

This paper cites ScissorBot: Learning Generalizable Scissor Skill for Paper Cutting via Simulation, Imitation, and Sim2Real.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation ScissorBot: Learning Generalizable Scissor Skill for Paper Cutting via Simulation, Imitation, and Sim2Real

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:03.745987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:03.745987Z digest=sha256:961e44e57c2f5a8f298f2012b70de943d872c5375873bf600f8b48382dff20b3

Observation 90c90874-ddf0-4f14-b451-50dcbd0a5662 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Open x-embodiment: Robotic learning datasets and rt-x models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:03.846413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:03.846413Z digest=sha256:540eac92a6d8677df07aee2b45c151fd1376bb279365340caf2f5bbbfe97271c

Observation 909cf3f7-a003-4757-8026-73c9a3ac6d0a · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Octo: An Open-Source Generalist Robot Policy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:03.950185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:03.950185Z digest=sha256:a5b0bdae33a1589e4a9c1a1a312460657f53e02e4939e68ce829c42b7c97c615

Observation 78ac52c3-b240-4701-b864-ccd4b8d6fff9 · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Unleashing large-scale video generative pre-training for visual robot manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.104267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.104267Z digest=sha256:88c2c022e54233396aa594bff67a517d71d6c9da89dab240b29393320e11d57f

Observation bc9485fa-6219-4cfe-866e-769d957b8ef4 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Cliport: What and where pathways for robotic manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.217696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.217696Z digest=sha256:8a0cd2fa1fb206b912933bdb1c3bc77492e734abae8184edadc69a9c1b00ec08

Observation f064819c-76c3-44ea-82c3-1336c4c75be5 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.300469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.300469Z digest=sha256:53c1b2816a98e351283e86701700e475adad9031b10f1518da3b81ad7af1fdff

Observation 5921040e-8d6d-4937-8a54-344361682922 · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.411902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.411902Z digest=sha256:a23a9c2db3aecdd6d3b23dad81e5f3872eb874906d0d300c87ffbf0eab5045a7

Observation 3fbf0f40-924e-4d9c-88a6-62798d7b7426 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation PaliGemma: A versatile 3B VLM for transfer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.500583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.500583Z digest=sha256:88cbe747ffe8a1d5d3d95843f874a742d497a5afde84c2984efc8d3911574009

Observation e8f7c67c-294d-43c4-abd9-0d4bf78fa1ef · outbound

This paper cites Vision-language foundation models as effective robot imitators.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Vision-language foundation models as effective robot imitators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.608311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.608311Z digest=sha256:9b1c235e14e91a62b5aa1ff251947b8da3f842ad82e93bf41c383136dd89a3aa

Observation 8923ce87-c679-4116-b4d4-04d5b2c5308e · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.718646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.718646Z digest=sha256:48b39ab4a9dd2cbe13e702bf34b028cece4976e5bd9799fc27cf66525e673725

Observation 0913954d-f737-4d18-a8d4-c57a5269c5d8 · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.781328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.781328Z digest=sha256:3edd39371d62e19dbb21166a18e4017dd2b499fb701de37fc328ecd61fdb7129

Observation 6098eb69-1345-472c-a857-7143b4121e82 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.850888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.850888Z digest=sha256:fac750b4491ece03b5f4343fdc1deaab5cd780951224e89435aea3c35fea4aca

Observation 87aaae01-72b2-4596-833b-9378e6f387a7 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.939146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.939146Z digest=sha256:22a2a3ce0488ea6ffa1aa6cfed932dc3ff337fa571e08bfdea7813ea58cb9fd6

Observation 084fe1bf-f199-42c7-be6f-d739ecef2c2b · outbound

This paper cites Flower: Democratizing generalist robot policies with efficient vision -language action flow policies.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Flower: Democratizing generalist robot policies with efficient vision -language action flow policies

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.943501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.943501Z digest=sha256:5e6c2787a0a82626608235a6f4f42fdd88ff14341b4310c071d9d09390969b69

Observation 606138cf-ce84-4313-9db3-340f266a5d47 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.947501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.947501Z digest=sha256:d6a3b195e0228d9a5d20c79c6939d56ccfcd79888f69658a5b33697b6f1cd987

Observation d4b24b7b-723d-4b18-8249-1e491be764bc · outbound

This paper cites Rt -trajectory: Robotic task generalization via hindsight trajectory sketches.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Rt -trajectory: Robotic task generalization via hindsight trajectory sketches

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.951644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.951644Z digest=sha256:a0bcff87ddbb535658e8937a6ce4026f319918555ab0e98f8a9620b1a1c5861f

Observation 8804f380-13be-4d91-b5d0-bca1c2acb3ab · outbound

This paper cites Pivot -r: Primitive -driven waypoint -aware world model for robotic manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Pivot -r: Primitive -driven waypoint -aware world model for robotic manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.955986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.955986Z digest=sha256:566fbe51ef794d1590bff434b3ab1dc7dacf481b442c520ef8febc689612bb7a

Observation 618d9953-7972-404f-886b-27fd379a01ee · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Any-point Trajectory Modeling for Policy Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.960054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.960054Z digest=sha256:c1d651cb0ce417f9df7b46d60fdce1460de66633cffc5ae1b976c3d13c4ac132

Observation 395e1886-863e-4193-a17a-3b0831c7a8b4 · outbound

This paper cites DreamGen: Unlocking Generalization in Robot Learning through Video World Models.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation DreamGen: Unlocking Generalization in Robot Learning through Video World Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.963693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.963693Z digest=sha256:b8d53f3821b033316434082ff373e0143db35c5ac85b318e3fbf991dbe3987eb

Observation 6ced7e81-5fed-4373-b3d5-fc5d535c9873 · outbound

This paper cites LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.967400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.967400Z digest=sha256:3a7fe8ccae84ff55ff0f0de3ff43a4dc22fac2ce272ac979bde07384910da2fb

Observation c8cd7914-b628-4f76-a057-a5d7f4368dd6 · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.971064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.971064Z digest=sha256:d6f7064b2e82886aafc125e555f331f240113f2cae75b76029b2a9886ecd323b

Observation 20150314-224b-4869-811b-611245e27024 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.975201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.975201Z digest=sha256:5f35cf73c4361c2e9ccbe81be5db7e283447ab12195cbc23edd474e7c103fbd6

Observation 321b86d1-5eed-46f6-a098-d73b5d0b554c · outbound

This paper cites ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.979225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.979225Z digest=sha256:c7be1b7516cfc152458f3c209a304cbbbaa3b68e1ac5a01d4dcd4ec62a13f030

Observation 94347f0e-5c55-40da-b219-b9dd290ba06a · outbound

This paper cites Shapellm: Universal 3d object understandi ng for embodied interaction.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Shapellm: Universal 3d object understandi ng for embodied interaction

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.982966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.982966Z digest=sha256:f390892d6d6a1d9ee95baa2bdd36d6270a9c7943c58766784126e22ad0d00948

Observation 98b084a0-3c45-48bc-b394-b6f44d0e959c · outbound

This paper cites Navid: Video -based vlm plans the next step for vision-and-language navigation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Navid: Video -based vlm plans the next step for vision-and-language navigation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.990360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.990360Z digest=sha256:bdd36cdea75d9196c65a5fdbafb364393ab07a329dd280f6be5310b01fc9fd9e

Observation 75f13b23-9f66-44ec-a973-67f421241755 · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.994004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.994004Z digest=sha256:242beaca1176d548c9fb589de0f617b2d60d300ca879dfcdede65d005ba446bd

Observation 0726e0ed-c6ae-45ef-9731-0cd6a1677d2e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.998053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.998053Z digest=sha256:d7e61e1b9d7fc570c511eb2bc973367e784896dbebb6669974d0ed488d262604

Observation 2fb56b77-2533-4c5a-8b65-606833ff27dd · outbound

This paper cites Openaio3ando4 -minisystem card, 2025.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Openaio3ando4 -minisystem card, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.001781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.001781Z digest=sha256:eaea340fcdd84202a5f94a3a44e06565b6f953740dc2a5610c0ea24b8e1ba504

Observation 9de293f5-89e8-4ded-81ed-12a26a464d49 · outbound

This paper cites DreamLLM: Synergistic multimodal comprehension and creation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation DreamLLM: Synergistic multimodal comprehension and creation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.005534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.005534Z digest=sha256:fb40ded1f97340d002542bac1eba98145f4fcdd43efb620c540c19603a078e40

Observation 01514cb3-0424-4a99-80d8-0ba8944d8162 · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.009025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.009025Z digest=sha256:ec66f1d878cec19a1ee95dee894a4e7ddfaf373490c60a44528d946a8d500830

Observation c84d68c7-2e32-440c-81b3-6844818041ff · outbound

This paper cites VGGT: Visual Geometry Grounded Transformer.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation VGGT: Visual Geometry Grounded Transformer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.012742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.012742Z digest=sha256:bb717875ced45ca859861f27b5e316608d351d092c553d40bccd3c5e9fead286

Observation 54b63741-d1b0-4da9-9223-423895f78d7c · outbound

This paper cites GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.016538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.016538Z digest=sha256:43d325abe674e05d17b75505b30aa6dc6be1e092b89d424f2955cb5ac0c0933b

Observation b7fd9e90-8197-4df9-aa5e-c31a549271e5 · outbound

This paper cites Tran, RaduSoricut, Anikait Singh, Jaspia r Singh, Pierre Sermanet, Pannag R.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Tran, RaduSoricut, Anikait Singh, Jaspia r Singh, Pierre Sermanet, Pannag R

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.020315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.020315Z digest=sha256:3c00464e74b6bd53ce2e406050199ea0682318b8311481262069b1297f0f7f01

Observation 774e5b71-a8ab-4319-bebe-d637b2c8a2f7 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.023640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.023640Z digest=sha256:96a4154229d63492f749275a26923523555c9c2ca5ae6856957fcd7cdc6b5877

Observation 203de78f-01e3-4723-a92f-d0d450987076 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.027150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.027150Z digest=sha256:880cac62b2b5840b37a90c60035bb8b6c7953e2a27756cb6b93380a15928ccda

Observation 150d728d-7652-4a3c-bc24-46f9e1ba84ac · outbound

This paper cites Hume: Introducing System-2 Thinking in Visual-Language-Action Model.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.030889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.030889Z digest=sha256:3c6b32763cf882adb12440d99b40daa5e85283cf44e09ba76897fbbd357a5962

Observation 6076b4c8-a8d6-499a-97b8-44d67fef01f6 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Diffusion policy: Visuomotor policy learning via action diffusion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.034752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.034752Z digest=sha256:510a4a32316afb0307c32460f7d28befef8e69c39c16040ae480315c9c545882

Observation 7f3888f9-e553-430c-aa38-42c8124f6990 · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.038317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.038317Z digest=sha256:473bc196571cb072cd41659b2041786bcf03e843374758f456d6b4644e22ced0

Observation 5f5aed65-ba22-42e8-810a-6e8f0cf1c86d · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.041939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.041939Z digest=sha256:7a3a9eca46fa39464406917c7960c515e126f8da381155cc0c9e2110dde1721b

Observation 1c77afcd-0230-4624-8cc2-b75bbbbdefe3 · outbound

This paper cites 3d diffusion policy.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation 3d diffusion policy

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.046023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.046023Z digest=sha256:911f2cd7fadf278674858953e77a5f163fce3a1be0684a0f7760dcfdaff81747

Observation c81b74bb-33a3-4b6b-8475-0a65ca3cdb35 · outbound

This paper cites Calvin: A benchmark for language -conditioned policy learnin g for long -horizon robot manipulation tasks.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Calvin: A benchmark for language -conditioned policy learnin g for long -horizon robot manipulation tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.049583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.049583Z digest=sha256:7aab4e41edf773a5f1a67e6640391aa02b16c8c2bba90059313b6195997b9532

Observation 3fa5d9a2-08b7-48f7-bc86-b4584ae0e0c8 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.052950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.052950Z digest=sha256:55c333f4f3cb7212649e2e91b065dca4600be9e70a9d0e665c5ce81291888e3e

Observation ed39de5d-5c00-446f-8612-24bf79080585 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.056655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.056655Z digest=sha256:c8225eac5f3788698de56143b3bc27b5feeef191b5bd8f6363088d80d9b4c155

Observation 13cf5b98-598f-4cab-8cfc-324181841f03 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.060322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.060322Z digest=sha256:5cef19cee7fd324e7752058d830eb56686aaa492c0b263438c70f281db63160b

Observation cc921bac-ba23-479b-bddc-a337264d2d39 · outbound

This paper cites Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.064158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.064158Z digest=sha256:fbc7eeee4f5420cad338bddf74083eb6e59d74862cae42e03e4afb8e3cce4e57

Observation 4fd9bc63-4a2a-4dd1-a853-da566f77e283 · outbound

This paper cites LIBERO: benchmarking knowledge transfer for lifelong robot learning.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation LIBERO: benchmarking knowledge transfer for lifelong robot learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.067919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.067919Z digest=sha256:912af50bca9f7bcd97abd98316bc82ce04e4b7c03da232e47ac96d10e085efe1

Observation ca160281-d00d-4ac7-878c-a2a8755c02bc · outbound

This paper cites Decoupled weight decay regularization.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Decoupled weight decay regularization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.071271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.071271Z digest=sha256:1c6b68f0fafd231511cd26252eb8486f0b09b9e350d1a499fe77b0a29e2cf076

Observation d9c0341f-9bb3-4bbb-ac8b-a4f62381251d · outbound

This paper cites OpenAI blog, 1(8):9, 2019.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation OpenAI blog, 1(8):9, 2019

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.074641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.074641Z digest=sha256:96d04a1d36c68ae9f6f27863a30951e6ffe647a792a80962fb2c2bc5d3c3934a

Observation 9be7d474-c644-42f4-bac5-c1cbc7faa9e4 · outbound

This paper cites an unresolved cited work.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Unresolved cited work

Reference 238

Resolution
parse uncertain
no resolver link, observed 2026-08-02T20:18:04.986810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.986810Z digest=sha256:224b8727181eb3180a6ad1f42ff4a8bf47320e1c6f53178e32e56796d8e19865

Pith citing papers

No inbound Pith citation observations are available.