Pith. sign in

Paper Citation Record · LEDGER

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

As of 11 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 11 inbound Pith citation observations for arXiv:2512.21970.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.21970 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:00:20.896966Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:19:48.331513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:33:45.452965Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46c1f415-e2de-4090-84b6-4c84f4b9f9bc · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision PaliGemma: A versatile 3B VLM for transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.737775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.737775Z digest=sha256:bcbdf2413813922939412c491c0f9113979a7c50e5f8df91c22409dc8b824ba0

Observation f57f7164-e780-41dd-b100-7eae2098dd6f · outbound

This paper cites Qwen2.5-VL Technical Report.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.741580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.741580Z digest=sha256:f653930e282a7b0b5f5fd8196cb1d0a05dc04a85f39efe049089f8444e4102fd

Observation fc92c517-4865-4452-bc7b-95c19d52efc7 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision OpenVLA: An Open-Source Vision-Language-Action Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.744666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.744666Z digest=sha256:0cbc868cbb53332835b7bec5b508f3d5f610cdfb9ecf17763798ae18793bdd87

Observation ba07debc-8ef5-439a-8209-c1b2b6bdb2a3 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.747597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.747597Z digest=sha256:30c7a6ca910aea8c15fa0cf8c868a63bb444bb3311461bf0e1b58e3905c2c38a

Observation f730debf-8187-427d-967b-d0a5525fe6cd · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.750637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.750637Z digest=sha256:0a2abfd33991d8342e13ca0fa55337b97cce85a3f3db0a5c2d1c3bfdd25d2abd

Observation 6825f712-3351-4e25-aa62-63d2296ebe0f · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.753591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.753591Z digest=sha256:d256c8873dcafccef816e52bfffa2a8f06e1185106881c213c51169a9626957e

Observation 03fd8ee7-6853-4824-ae1f-c8df13e08116 · outbound

This paper cites Rvt: Robotic view transformer for 3d object manipulation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Rvt: Robotic view transformer for 3d object manipulation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.756756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.756756Z digest=sha256:55f6ef6dac5e2c7be3bfc1e8e747d517f0ef7de9c2cb2bd9b55d8492c8e7c69a

Observation 662f0ae5-f798-498d-9708-7c14a6018916 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.759135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.759135Z digest=sha256:4221f660cfb4eaef8a195cdabc5b11d03f3a1f59d001338f42471f5f12e7a791

Observation d0453c9e-7fcb-4835-9a80-0b48797b0afe · outbound

This paper cites GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.761836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.761836Z digest=sha256:e89ca8f602f0195f41b976072c160603398b07dbbc4665868ec582e5f86a5129

Observation 6c12eae7-0785-4216-8443-8ed0825365d4 · outbound

This paper cites Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.764446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.764446Z digest=sha256:613772f6ee46e4509cd132c865cd0f4e2f11eee811e943efaaad1cb5038b5604

Observation 6e4c4c3a-4b9b-426e-968a-47eabbdc959b · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.767550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.767550Z digest=sha256:a6c2f9650440ef0ad133d0c90986969972428a3d721a33fef3834560e2d6e7fc

Observation 181423dd-04dc-410a-99c6-a4bc3b9677ba · outbound

This paper cites Foundationstereo: Zero-shot stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Foundationstereo: Zero-shot stereo matching,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.770358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.770358Z digest=sha256:9229b07b087e3be3ef9e76444d527a0d311ac3b631db3d9a556f1c99a3c04a58

Observation 554549d3-8d71-4939-8218-cd4de7077f43 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually- conditioned language models,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Prismatic vlms: Investigating the design space of visually- conditioned language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.772653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.772653Z digest=sha256:7a5070d3e04e3cc1559f21a975d09d1faa307fa2eecdae6686b9a79605cf5b81

Observation 39029919-013e-4d46-853b-efa44f1a2299 · outbound

This paper cites Palm-e: An embodied multimodal language model,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Palm-e: An embodied multimodal language model,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.775125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.775125Z digest=sha256:f4eeecc38323222669054edc2a0ff5d17f688279c5cc771e90eb4509b43e9cfa

Observation 4f73e351-04a5-481c-834b-d809dec0f86a · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.777481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.777481Z digest=sha256:0fe486c42c8081302018036de79f56d5c2be473ddfad6eee099123760787837b

Observation b43e69d0-a3c6-4f6e-aa8b-bb178e9ae143 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.780838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.780838Z digest=sha256:8b1aa09d05cc86efe8ac6cf44ad1904f93de64ff6060779ea95acc0791ecd083

Observation 3a851f77-53a9-486b-bb2e-7e3a89c5e85a · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Florence-2: Advancing a unified representation for a variety of vision tasks,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.784051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.784051Z digest=sha256:cde489eee5d43d9d2330469174e93fbba4ffcd297447dbd6bd92e754d073b18f

Observation 382014e7-f9dc-47f3-9831-1d1944bc0ead · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.788875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.788875Z digest=sha256:0934cddaf52c69c02e7bf4485fa0aa99d705bec28128a95cf249778f53ffbd4b

Observation 44ed9fe0-4171-4711-b7d2-4fd74768c3e7 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.791509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.791509Z digest=sha256:1247e05db019051d384c169dc15b1fa0dae9a861f8bef3ad0f3ae3486a6350a8

Observation fea9d4a4-1c36-40ec-9b5f-a6275e93d73d · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.794026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.794026Z digest=sha256:1be094c5d0b1e5634ec3b0bc6d12e810fe119e624a7bf9003278f841e1cede9c

Observation eca535a6-128e-4486-811c-d410d44d37d1 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.796749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.796749Z digest=sha256:1eb5cc75cd6e71a11120716cafd1d6f5dfccae50370bf416584dd605d4a51101

Observation f056f3fc-099d-46af-bbbc-0f95a73d027b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.799400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.799400Z digest=sha256:85c7da921a57a7043f4527b69f5531fcdca772354b7b8aa85b08faf4327bad56

Observation a8884fcb-8c89-4f58-b5b6-ee2dd654999b · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.802389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.802389Z digest=sha256:930b1cc956118a7da71ff33985a386c1b3e4591a15506812e232d005578eaac8

Observation 6fbd01ba-5fc9-444c-8fa7-850c96c17140 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.804995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.804995Z digest=sha256:77be36e2689f68cb45bcae0ca1e8568689afc078f261a08c80abe61559192602

Observation b00390d9-18ec-48ce-a580-866cfea3d4fd · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.807558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.807558Z digest=sha256:6d09dbe2a841a2704042d9c3b30e9c240ce56fa56463f6c03283513fa3428845

Observation 97dccf5f-d256-472f-b89a-48a53538c6ff · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.810263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.810263Z digest=sha256:a4427fef719e340cb39dd9a931bfdb5494b91df60cba6c0f4a85a96b85b35044

Observation f0450714-5eae-4d4f-8f1f-276e0e82d94a · outbound

This paper cites Internvla-m1: Latent spatial grounding for instruction-following robotic manipulation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Internvla-m1: Latent spatial grounding for instruction-following robotic manipulation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.812858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.812858Z digest=sha256:f01c2ecd0fe0a3b4901c1fe001cac9bc31fc47101d9976a8b23ff2c26f399ef0

Observation bd8ee129-bbb4-4602-bcb2-e0b2d8c276ad · outbound

This paper cites Tinyvla: Toward fast, data-efficient vision-language-action models for robotic manipulation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Tinyvla: Toward fast, data-efficient vision-language-action models for robotic manipulation,

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-03T14:00:20.815149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.815149Z digest=sha256:28349089ad723147b5dd3c8ef1fb575e0852723edf758facb9f55c14df0e5f85

Observation 0c17244c-5f64-4dec-9b5d-a7bd080a0504 · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.817485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.817485Z digest=sha256:db65754de789ed95919b1a7fd25fd1f0066ce5ed2ebeb15a7d124fae04dc3dc8

Observation fbe30e33-4bb9-4f8f-a2f5-80d8c2b5d870 · outbound

This paper cites Generalizable Humanoid Manipulation with 3D Diffusion Policies.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Generalizable Humanoid Manipulation with 3D Diffusion Policies

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.820085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.820085Z digest=sha256:9179ee8de6c1c6ad9697e17cc5c21e39561d477e891dfdb78dbf0dc77bb89bb1

Observation 60ae4b1e-e783-49a8-8fe4-ca8fa6d8090e · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.822651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.822651Z digest=sha256:e67443bdf0453584a2e00aa52919b85a794d9c5c77314c727950f18426eddfa4

Observation 68888964-eab6-4441-904e-4ebba7a9d12e · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.825243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.825243Z digest=sha256:2c6b6267c1ff2d5e97408b4f5244378c27d833cc91187554fecd0809823a7b7f

Observation 3f516bb1-7f71-487f-a83a-b3ad062311da · outbound

This paper cites Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.827780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.827780Z digest=sha256:07c81a4f3557c00491e07f76e7ccd6707ae0e22e56f2eb1f15672bca4a6712bd

Observation 08b01426-c2c5-4119-a374-e7fc6457bbae · outbound

This paper cites FP3: A 3D Foundation Policy for Robotic Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision FP3: A 3D Foundation Policy for Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.830245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.830245Z digest=sha256:be2d2cbf07839061506f230cdc7dfe0f80878570cc91a483bde02daa72bd0a58

Observation f3e85bbd-a1e7-4674-b97b-17ba4355c0ab · outbound

This paper cites Evo-0: Vision- language-action model with implicit spatial understanding,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Evo-0: Vision- language-action model with implicit spatial understanding,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.832873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.832873Z digest=sha256:93585c6763d6342514889e97ab318bbf482547d6d4670d09974a89b516caa009

Observation 4b14c6b1-069a-4c3f-8d63-d0608475797d · outbound

This paper cites Gp3: A 3d geometry-aware policy with multi-view images for robotic manipulation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Gp3: A 3d geometry-aware policy with multi-view images for robotic manipulation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.835316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.835316Z digest=sha256:feffa159b3375a42a039460ebb0f2f0755c74dd3ee578c194833078de2759894

Observation 647f1a5d-5aac-435e-a2c3-c5b1b0cdfba5 · outbound

This paper cites Learning the distribution of er- rors in stereo matching for joint disparity and uncertainty estimation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Learning the distribution of er- rors in stereo matching for joint disparity and uncertainty estimation,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.837678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.837678Z digest=sha256:c8421ba669a7dec7b0c17b26e244993b992287762e40596db006fde7201ab700

Observation f295d6cf-0cf1-4c91-9e22-e1c45540f400 · outbound

This paper cites Cfnet: Cascade and fused cost volume for robust stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Cfnet: Cascade and fused cost volume for robust stereo matching,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.840251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.840251Z digest=sha256:a372e5ff99fed14c1105ed714e3efd65481e9afc24c7e3af3109e13e53f4eca8

Observation af5d338b-0d2c-4dde-94ff-729c23ab1b40 · outbound

This paper cites Pcw-net: Pyramid combination and warping cost volume for stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Pcw-net: Pyramid combination and warping cost volume for stereo matching,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.842621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.842621Z digest=sha256:3c567f6cf13cb9530fc89aaa57a1c807d6a6122fb6f862293a9e6cdf503f5a93

Observation 0be59683-b8ed-47d8-b656-0a7af496d967 · outbound

This paper cites Aanet: Adaptive aggregation network for efficient stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Aanet: Adaptive aggregation network for efficient stereo matching,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.845047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.845047Z digest=sha256:221ab501e52515997675d5f84616665c0648c1a1ade4d11d7498d07bf4c6c318

Observation 689ded80-bfb5-4e9a-83a1-2fbaaed4be91 · outbound

This paper cites Arunet: Advancing real-time stereo matching for robotic perception on edge devices,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Arunet: Advancing real-time stereo matching for robotic perception on edge devices,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.847397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.847397Z digest=sha256:fd6d6f7a59b4d5aa06f32dcd2c6c4ef90189d47f6f9273683ebccd4d4e093166

Observation e4f029ab-af3e-4d83-b1b9-4943b8887eac · outbound

This paper cites Gfanet: Group fusion aggregation network for real time stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Gfanet: Group fusion aggregation network for real time stereo matching,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.849923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.849923Z digest=sha256:3caa5981fd324ab9985de4dde1b1e5d353ade51e50da4d6f100d42e201623c66

Observation 40a17586-476b-42e5-9d95-d701b592fa3d · outbound

This paper cites Raft-stereo: Multilevel recurrent field transforms for stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Raft-stereo: Multilevel recurrent field transforms for stereo matching,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.852639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.852639Z digest=sha256:229d63392290ccb1233ae91b7c65948a140335e698463b7ed1f52c84ef50e4bf

Observation b657e44e-d0cd-407a-8b97-780df2ec3e1d · outbound

This paper cites Iterative geometry encoding volume for stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Iterative geometry encoding volume for stereo matching,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.855173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.855173Z digest=sha256:2241fe8ac38d3b43a0f4af4f8b94523e8aaf5ea401f84af8069e74fe44d3de02

Observation 821d3cfa-00e1-47ac-85f8-27b550f0867b · outbound

This paper cites Practical stereo matching via cascaded recurrent network with adaptive correlation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Practical stereo matching via cascaded recurrent network with adaptive correlation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.857433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.857433Z digest=sha256:d249f593644891bb42de397643ccd3118a299c65877773ccfc914eacdb83e021

Observation 9cb26f67-0881-472b-849d-e25512166d71 · outbound

This paper cites Uncertainty guided adaptive warping for robust and efficient stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Uncertainty guided adaptive warping for robust and efficient stereo matching,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.859742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.859742Z digest=sha256:2c171c84bb865e892e747e58e540219e8af428abaacb2052bd9eedee90bfefb3

Observation f7d67585-2b8e-409a-8f2f-255dc443fb49 · outbound

This paper cites Learning intra- view and cross-view geometric knowledge for stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Learning intra- view and cross-view geometric knowledge for stereo matching,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.862334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.862334Z digest=sha256:76eb4a83f5d7dbb56a3770a3bae9d3fb936573a4b0240a50b8d3b74e6fa26e7e

Observation 536476e5-9dc8-4179-ac95-322f49b11667 · outbound

This paper cites Efficient and hardware-friendly online adaptation for deep stereo depth estimation on embedded robots,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Efficient and hardware-friendly online adaptation for deep stereo depth estimation on embedded robots,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.864728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.864728Z digest=sha256:39c5234898e6a80d61c24cf6a136a8a2d404ed2a03152c7b68f7d503712d34c1

Observation 68fd4d0a-e8fc-4911-b6fa-302c7204bdc8 · outbound

This paper cites Stereo image- based visual servoing towards feature-based grasping,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Stereo image- based visual servoing towards feature-based grasping,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.867054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.867054Z digest=sha256:2137e673e874db361edd740ec994549fc26a922a2294163c8239be75f7d7b172

Observation 6c24e98b-43b4-43c6-a2d6-b31742e32627 · outbound

This paper cites DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.869541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.869541Z digest=sha256:caaa7e32531047a8f6d21b9f87335e57d2f8ee9d1c6b734be330c0a55bf1189f

Observation 52159cb6-e386-4419-bbfb-962ed03dc739 · outbound

This paper cites Simnet: Enabling robust unknown object manipulation from pure synthetic data via stereo,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Simnet: Enabling robust unknown object manipulation from pure synthetic data via stereo,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.872190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.872190Z digest=sha256:4a20bc302c19c68ff2f08f61af5928fe87ed072bfc34af9ac38c248957decfb7

Observation 183fd8e0-b25a-4fcc-a077-b286aba51e48 · outbound

This paper cites InternLM2 Technical Report.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision InternLM2 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.874567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.874567Z digest=sha256:a87f3df1084779d0b2d2381b50ec820319bdf49dcc7635fee663fd7400e33915

Observation 9a010d3a-800f-4e66-b2ae-51f5b95e2bb2 · outbound

This paper cites Flow Matching for Generative Modeling.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Flow Matching for Generative Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.877040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.877040Z digest=sha256:8b3df6158f24da817b69aa7c6a9cf02712c4617c55c62a584155d4bf7c408938

Observation c60bafb7-2202-42ce-a5fd-18448b6c4d0f · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DINOv2: Learning Robust Visual Features without Supervision

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.879627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.879627Z digest=sha256:bd4f7ab0ea4f0aeb7573261ef65991afa12362725dfc399b35be7d781ba8ed27

Observation 08702300-1375-4a36-89de-93277c83f09b · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.882762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.882762Z digest=sha256:6b778a9dfa88e9dcd475fde437304ceee1ae9d49ba14c0415c0a3649baa39dfe

Observation c7b73b3f-2007-4ada-b7e0-853868a86f79 · outbound

This paper cites Mujoco: A physics engine for model-based control,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Mujoco: A physics engine for model-based control,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.885441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.885441Z digest=sha256:c08fe63013f5955809a45468ad30684ae7c959b425567a6f3cf4d6caa35d3e52

Observation a3143be9-2046-47ad-8b82-d46d4be7beb8 · outbound

This paper cites Isaac Sim.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Isaac Sim

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.888280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.888280Z digest=sha256:5e5d8e763e884d932091f2cb3cd8a8a67ca931b9b1d1db1c22f699c849395adc

Observation bd27a75e-e7af-41b7-9a7c-8dc80099b7c9 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.891559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.891559Z digest=sha256:d7494ebc5ba38060fa16215214b51493d6df32cf6e0321fb60d68f73cc4ca78d

Observation bf7fb575-bf4f-4a41-918c-8fbd10c5ec33 · outbound

This paper cites Vggt: Visual geometry grounded transformer,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Vggt: Visual geometry grounded transformer,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.894490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.894490Z digest=sha256:085abfdb4a0876d57d4a0a255f568ede13f543173cdb0e0b45d6e070e4f08537

Observation 0f44e8f0-4a64-4c6a-b38c-f743f13a6696 · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep network,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Depth map prediction from a single image using a multi-scale deep network,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.896966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.896966Z digest=sha256:7b5a41ce0945867534dd10b079fe1d2e3c1624f0f0dff3be9e651c06e05e4a45

Pith citing papers

Observation 29bcc586-fdd6-4570-a3f1-1810b5506305 · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:4183f4e9d19d11e545bc414764628b317067c8cbfce1677756e87b5a4e9ae0d5

Observation ff9f8454-ac41-4d72-a36f-9a1ebc423a2c · inbound

MolmoAct2: Action Reasoning Models for Real-world Deployment cites this paper.

MolmoAct2: Action Reasoning Models for Real-world Deployment StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T17:53:44.901684Z digest=sha256:2d87e0ce417571b95cddf21eece43131e9825d53834f0d844c6592f8ef54042f

Observation 5f5dc25b-ad35-4b5c-8263-6d52775a873b · inbound

MolmoAct2: Action Reasoning Models for Real-world Deployment cites this paper.

MolmoAct2: Action Reasoning Models for Real-world Deployment StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:479b83d9d22bf7f9f99c897b6141d5a425bda93fb1843c7cbfc06d02470cb21b

Observation f801e2ea-02e8-4226-a404-f3f954eb394a · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T03:57:10.338617Z digest=sha256:fb7ce34bab7ab0e73956f48c48cf31c2d1f6ffc1473f1bea09da42c966999da9

Observation e57183a7-40e2-4b37-981a-ad9601eefce0 · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.683779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:8c332aaa25f38cbb8e9e61085790d9193ee17f6a6d854d40fba37fdf5a006673

Observation a7561048-f502-4d09-aaf8-2992e5b5d62a · inbound

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model cites this paper.

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.196481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T21:39:15.662180Z digest=sha256:eaeffe6004135100bfc4ba04e19597d578c6e6b54eccf9ff87924fa127da9fb6

Observation d484f26f-42dd-4851-aad3-6662d8f9ab1b · inbound

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning cites this paper.

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:20.378897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T14:41:50.084254Z digest=sha256:090446bb38aa4e5d2881905b32113010a91d60439478eb4133ba89fd66b9c15f

Observation a7e3c69d-5d32-4bf0-8d84-b30f4260f15e · inbound

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination cites this paper.

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:37:37.072871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T13:47:28.832078Z digest=sha256:b67e1a44cf270c5d0f2850daea4a168e56f63a4914ad599762e2bf6ae4df7399

Observation b1ebb169-2504-4e9d-ba44-7e4680adbf17 · inbound

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model cites this paper.

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:04:21.054350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T07:01:48.330369Z digest=sha256:a441f1d6c0ac73d325dc7d93f567668223768fb8454c7b8cd7805c2c169d20fa

Observation edf36261-3759-4659-977e-f093520ed7b4 · inbound

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model cites this paper.

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:33:45.454775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-07T12:28:21.318505Z digest=sha256:2a13eae58f06539717c0e72748c1e94a63d11619aa4bb99f2743114d90a23533

Observation 752060ec-438d-4d3c-ad9e-06c242fde276 · inbound

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models cites this paper.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.331513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.331513Z digest=sha256:352428b8e5aea93dea2feb024aeb26e99505bb59aeac212d439139998d2da711