Pith. sign in

Paper Citation Record · LEDGER

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2608.04633.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04633 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:19:50.356389Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26f9b6b8-c9c3-4892-ab74-7105435e5b9f · outbound

This paper cites Zitkovich, T.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Zitkovich, T

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.396151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.396151Z digest=sha256:0019a826e7a685c32ac760cca15c70803dedf6b29b5389638226b06918a7538a

Observation 3f472e05-b48c-4833-8843-c2850544695d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.446123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.446123Z digest=sha256:4888e4c9a13871ffca3c9fa8144dcbe55440561ed9c24d742a1620825331de51

Observation b922b83f-7d15-4061-8911-9e7cd6d05b4e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.520992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.520992Z digest=sha256:8a861d5c4866349aff7ddcebebe08af4a58b22977f99fa0515530c46469e254e

Observation e68b6461-d7a1-4051-8dcd-6f2b3aea963b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.577353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.577353Z digest=sha256:a2a85cfde3b7c7fcdf3420fd93b0955a96a9a062f4f50b6427fc0b3e973c8bc3

Observation 5194fd60-de51-48a8-aec8-e97196c28576 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.863591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:19:46.671171Z digest=sha256:54a3db4c22eb343fbcdd8550c1c75633c3bc65d0364af57a4746d47f735ed7b4

Observation db459a88-3610-46fd-9aaa-6187dbaf2234 · outbound

This paper cites GeoVLA: Empowering 3D Representations in Vision-Language-Action Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.803171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.803171Z digest=sha256:87b6ea888936c9dc6068321588acab2d14cc3835ad6e48924eb07952dff3c49b

Observation 9cd5a458-dec3-4ee2-b750-754efc5b8c04 · outbound

This paper cites Singh, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Singh, A

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.923578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.923578Z digest=sha256:1d6514a7844252bce4ab13cf3f33fc0add7d6f747408c3aae416efb82608ad7a

Observation 92770e72-495c-422f-923d-8c5a48bad184 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.003040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.003040Z digest=sha256:84d1518bd1c132a24bea3f161512d4fbf6a4794f2efd4bd3cd4e4a9072ea9d58

Observation 13e9efc7-a1bd-46c6-8d49-60b75febcc95 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.078147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.078147Z digest=sha256:19ac18947a264fe928b8ec5f9173919231ed4f04bea4685a437d2d3a9517535e

Observation ab78dc4e-908e-443f-8b81-016d8a0891fd · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.636357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:19:47.155194Z digest=sha256:337414b3edab50969bbb14782e262e68693237b451b12ab1a8be63aa93433214

Observation 1e5c079e-cc74-487e-944a-c312cd13ca68 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.208817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.208817Z digest=sha256:0f10d05e3dde5c3cd9cf8ac482395688264d1b3e64208eb37f95f5ec2125a332

Observation 1f34d831-2310-4095-a423-78be06bfecb4 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.275991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.275991Z digest=sha256:9d4e8347742ba7eee3a64e1d312d4182228f542a7e33fc940970b3fa2ca1c163

Observation 677ce944-fbae-4bbe-babe-38852f0f9f5e · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.408288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:19:47.359540Z digest=sha256:9444c632b19f7c58a838213b3965e1c9f9d0c5f47c4a25d4f7e3330215be4f3c

Observation 6238bce0-fabf-42d2-a31d-6f4a6730eef1 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.403646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.403646Z digest=sha256:222e500c80326788010d63ed3aece59991e9ec788ab236107b460f158e14869d

Observation def2df4a-59f4-4fa3-9f90-ec4a4a19695f · outbound

This paper cites O’Neill, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models O’Neill, A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:19:52.215031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:19:47.510694Z digest=sha256:152bf96e173a120a708093bab5c83c81f2043269dd0b038cb6c15e295cf09a79

Observation 567b9100-3b8f-489b-9b99-5978f5f252dd · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.572949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.572949Z digest=sha256:798fec8ac99d70cbc4de796589212c46ec8010820e90010cccbf60962a8f7c39

Observation a909a4b9-f254-43ae-83a7-e22beb324557 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.653884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.653884Z digest=sha256:12a037af4701940e83494e188c67a6621bdbf226d6b8c69886131fbeee7e809d

Observation 29970e42-475e-44e3-9b81-74697dd96599 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.739039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.739039Z digest=sha256:b7c8bc160625f0c077b7b33fae3a8d8458ea937ad432fb3ebbe42129915e6d86

Observation 22e3fb0a-6958-4a39-a54f-fa6f9591b62a · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.832939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.832939Z digest=sha256:7878c544767add1febff4f4842de651af9ff61cb72c9209c0298a5183be3e125

Observation d55763d2-e79f-4c5a-aa60-919c89b1877a · outbound

This paper cites Shridhar, L.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Shridhar, L

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.896250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.896250Z digest=sha256:3b4efa25319faaf0a5e1de5d014bfa5e72b867855450325a4af0bee893936a97

Observation 5947d0b1-08dc-443b-a0ca-54017e52ee19 · outbound

This paper cites Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.935248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.935248Z digest=sha256:2bd2ab72423631d385e0f41c2c765ae6a158f6a2f3d953c9bdc7606cb459b6bc

Observation 43fa7b85-4d46-44ac-9aa4-378678dda7dc · outbound

This paper cites Goyal, J.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Goyal, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.011334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.011334Z digest=sha256:2f704813484f695a65704ce124e14decdebccb00c98c51fc6fd5f2701cdaef18

Observation 3da0a491-2e47-4f1f-95f8-48f090164ce9 · outbound

This paper cites RVT-2: Learning Precise Manipulation from Few Demonstrations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models RVT-2: Learning Precise Manipulation from Few Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.095191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.095191Z digest=sha256:ba85db0293e289366410b25fa9d8d786760a04aab1d5c89de7dc37cd232cce53

Observation e1e07334-0fca-48ef-b91e-2331ec311f7d · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.187917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.187917Z digest=sha256:d1b6917d451bcea5a8f46513e239995004dcfbf92c65e3a865d6dfe6f2a27541

Observation b1fd0a23-d6f6-41c7-a9e3-d7f7e487a763 · outbound

This paper cites Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.268018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.268018Z digest=sha256:59e0526e30c9281934f7d169a5e0fe550ba82a96bd81fd9ef671bd1225c2c197

Observation 752060ec-438d-4d3c-ad9e-06c242fde276 · outbound

This paper cites StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.331513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.331513Z digest=sha256:f22476bbb5591a225bb6ca062584bcceba46cf1e6a088474afca36ce197e529e

Observation c046748c-8091-4aee-bbdd-93269d49b7a9 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.983253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:19:48.393864Z digest=sha256:bfd2cf4e324e4fc1255c3b7d8e4d5c253667f1563408f430c22ea0f3f182d0ba

Observation cc182fc0-0a71-483b-a2dc-10046dc23220 · outbound

This paper cites QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.478639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.478639Z digest=sha256:4a0a8ba1aa6bb7bc4f66fa8a8d5bf48500a0089aedf0cae319672876c44bd6a0

Observation 25f249e9-f40e-49e6-b0ef-f99092c85327 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.534851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.534851Z digest=sha256:17c5456df9d2877ed2ab4c76394b332de4e237da7514ce783923a388326adf56

Observation 767aec74-c80d-4bca-8b52-692215445dc4 · outbound

This paper cites PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.590609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.590609Z digest=sha256:24f12c6a21929ad99995d1a59d562fdcd03f50e5a2ab6c7804579eae1e0e0603

Observation a8669714-812a-40b7-8883-a0c45c288110 · outbound

This paper cites Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.667454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.667454Z digest=sha256:50372b2171b55e71f4d3810ecd800a3b9d2332e0b790fa61caca20d9e0f15d50

Observation cd9d14f8-9612-487f-8ec8-2cde133f0f58 · outbound

This paper cites Peebles and S.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Peebles and S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.767277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.767277Z digest=sha256:18d9a6c7e38635113955c7f5ba94f0c9594abda82fbfd282f094b1aae68aa2f8

Observation 06563ac3-47a8-4429-baa6-b30e995206c1 · outbound

This paper cites Radford, J.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Radford, J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.843593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.843593Z digest=sha256:997856fab686019a617b8689188459d1da9c895ee497f422a62294b3409b3cbe

Observation d672b3e2-c805-4b58-81a8-5925afb5e535 · outbound

This paper cites Denoising Diffusion Implicit Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Denoising Diffusion Implicit Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.975765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.975765Z digest=sha256:10542001df25168d2b6eb8aad12e076354e81593300e580211c6e1349364d252

Observation cc9f1646-a4c2-418a-b34c-61d6444f36c1 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.097273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.097273Z digest=sha256:5fa1362bf85cb375feae5d55c078be10cfc355d9005764664b9dcbd0031491d4

Observation 23102af3-4a0e-486f-b583-ba540c05c04d · outbound

This paper cites Rombach, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Rombach, A

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:19:51.738404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:19:49.191361Z digest=sha256:93734e4c891dad1fedd6dbbed3b089256a1d88b3379849546aa83050f6554a2b

Observation 1c149296-e834-440a-89e7-1cfb36079b83 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.316462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.316462Z digest=sha256:c6c4ed09751412fc2fd1b3eaa895998efbe25be81a0478e63d09e9d40e68d29b

Observation 1a4a7068-291a-4154-a9ff-248e47bae85a · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.424569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.424569Z digest=sha256:e5b789373da8edf3b4b0e647c5d690f51a3cc7bcadac59da25258a978c3ad663

Observation 37bcff39-8ec9-46f8-838f-803038cf6f9b · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.501211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:19:49.557016Z digest=sha256:986548c3d67e5b47db600891e24076d3bf1d33c56b505c5f6b531541cf87dffa

Observation d46dfc98-2a71-4997-be37-2ba63028c1af · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.656649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.656649Z digest=sha256:c9bb33507a3f6dbfc0dd6fe49161c70895fd15503f8782d1107f9576f4da90ab

Observation ef3bd510-365b-4b64-a3ab-bce873a810f5 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.758226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.758226Z digest=sha256:852446bc2cf031853db282b42ee0a28a04d00aad1640d4d391f39d14a71af117

Observation c004a486-b24f-41c0-8c5f-aa26d8d09415 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.893880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.893880Z digest=sha256:f9482da083089a8e780940ac00ce8474077933bdcd4c259d46791e77f114761f

Observation 58340e5a-1563-48b3-8529-a965af4e892c · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.283184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:19:49.988923Z digest=sha256:e0c4a51724c72897cd2e2ed0f7e72e9c2a8ee6e6063a21ae0ea77b0354ef3850

Observation 479ab36c-5a63-4a51-90f5-cdbe56196a57 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.092976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.092976Z digest=sha256:0a9d8fc8db91959d271f9242fca42af38b10610b3e57589cb25b8e9569e7ff8e

Observation 6e930a4c-d7e0-4273-b5c3-18d9d8fce342 · outbound

This paper cites Kirillov, E.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Kirillov, E

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.211537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.211537Z digest=sha256:4b600ec11b4035f0ec3a71c57a0de23e7cb86bd521afe598dd0f926edb3e895c

Observation c57b9e74-d5e7-4767-8b1f-30d0aa36d0b2 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.356389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.356389Z digest=sha256:f03849a82ad3e2d8f6d286c736ef5c0f38398748e68983471eacfa27ea993b14

Pith citing papers

No inbound Pith citation observations are available.