Pith. sign in

Paper Citation Record · LEDGER

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

As of 14 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2608.04633.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04633 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:19:50.356389Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26f9b6b8-c9c3-4892-ab74-7105435e5b9f · outbound

This paper cites Zitkovich, T.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Zitkovich, T

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.396151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.396151Z digest=sha256:c4ce2a5de5d31a2fbc363c814edadaa44c3684ff189334180e2e4d773ec72b4f

Observation 3f472e05-b48c-4833-8843-c2850544695d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.446123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.446123Z digest=sha256:f70058bee07a0d170dde531bebdd7247b4ed254f1646207f2c175c17d5cb4426

Observation b922b83f-7d15-4061-8911-9e7cd6d05b4e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.520992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.520992Z digest=sha256:2cc4a1a0909e79b735ff9a06ee0c92302faf9c96b4270bb4f57b31731f74d937

Observation e68b6461-d7a1-4051-8dcd-6f2b3aea963b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.577353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.577353Z digest=sha256:90af366c5c1d70bb1bdc72257bb234e101fafd409d7aac53002a154adecd73b5

Observation 5194fd60-de51-48a8-aec8-e97196c28576 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.863591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:19:46.671171Z digest=sha256:c50291f8bc6e64b8c385c68a3cbc5df988630f4286d3b5fca84a02078b004a2c

Observation db459a88-3610-46fd-9aaa-6187dbaf2234 · outbound

This paper cites GeoVLA: Empowering 3D Representations in Vision-Language-Action Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.803171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.803171Z digest=sha256:8bd2050f13cb0a11119d346746a43f01ec5abebeef922bd1cd28f24e1226e195

Observation 9cd5a458-dec3-4ee2-b750-754efc5b8c04 · outbound

This paper cites Singh, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Singh, A

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.923578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.923578Z digest=sha256:557af1d2995a3edcf611d1d3bf44b03b676be2b3e153fc7a84910d70abcb2a68

Observation 92770e72-495c-422f-923d-8c5a48bad184 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.003040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.003040Z digest=sha256:2dc866ae1a494a9b936ec2f02b4cbd7005a3c8af889368b250385c690e205e0d

Observation 13e9efc7-a1bd-46c6-8d49-60b75febcc95 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.078147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.078147Z digest=sha256:3c4109a0540a60f606eb78c4f1e4699b5da0aa54048a66e680b74dee39eafda4

Observation ab78dc4e-908e-443f-8b81-016d8a0891fd · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.636357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:19:47.155194Z digest=sha256:112f92c734e76351f3b4312bf41ba9651184a57a17b8d68ef324dd15544dad51

Observation 1e5c079e-cc74-487e-944a-c312cd13ca68 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.208817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.208817Z digest=sha256:994f130552986550a786f7a7758b2a032a644337ff638c5f27067a39ddd22b23

Observation 1f34d831-2310-4095-a423-78be06bfecb4 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.275991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.275991Z digest=sha256:04dc5bf120ac548223415bfe2fa1387debb841dbe5c1b97e1344125591fb5f58

Observation 677ce944-fbae-4bbe-babe-38852f0f9f5e · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:52.408288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:19:47.359540Z digest=sha256:fb1a9cd333d303eed4fb2ddfa411b953148cf95c5d73950f76c807c3cf78aa36

Observation 6238bce0-fabf-42d2-a31d-6f4a6730eef1 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.403646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.403646Z digest=sha256:b57781c96c1bff1092134f82497b939ad4fc2a40399c80411826b9fd2e976830

Observation def2df4a-59f4-4fa3-9f90-ec4a4a19695f · outbound

This paper cites O’Neill, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models O’Neill, A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:19:52.215031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:19:47.510694Z digest=sha256:49447969b657fd9928052c17e15d3e0e48166bf38f0733afe1fcc3d3abbc8930

Observation 567b9100-3b8f-489b-9b99-5978f5f252dd · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.572949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.572949Z digest=sha256:10d40e596e29a99f58c00b6b572b894f8c6cd5327e51876c3b612fcf091f3cfd

Observation a909a4b9-f254-43ae-83a7-e22beb324557 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.653884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.653884Z digest=sha256:2ba9f3ef575b0c0f69093f0aba953a6883b0234a0f5a66fa50bb738cec24ade6

Observation 29970e42-475e-44e3-9b81-74697dd96599 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.739039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.739039Z digest=sha256:77e926e0a9b2f2c3f8fea726432f07dd3bc0dd89b33f31524b9724afd4751d14

Observation 22e3fb0a-6958-4a39-a54f-fa6f9591b62a · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.832939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.832939Z digest=sha256:264c887bbd8dffdab0f73ef7f3c6ae74750b95fcc9378679fd965a994811e8c9

Observation d55763d2-e79f-4c5a-aa60-919c89b1877a · outbound

This paper cites Shridhar, L.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Shridhar, L

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.896250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.896250Z digest=sha256:dda6caf037c4a949c49037e03bb5ad88adcc5c0f094ec79540fcb132563aac3e

Observation 5947d0b1-08dc-443b-a0ca-54017e52ee19 · outbound

This paper cites Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:47.935248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:47.935248Z digest=sha256:f441dc3a8b2455e9ace1564f6442b8bc170802e232b96c9f26d8f1af1ab2a432

Observation 43fa7b85-4d46-44ac-9aa4-378678dda7dc · outbound

This paper cites Goyal, J.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Goyal, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.011334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.011334Z digest=sha256:2a1752794c6dc5890874f89c800aba84c5da0c70e9a31a66e0dfd0bc545fc63a

Observation 3da0a491-2e47-4f1f-95f8-48f090164ce9 · outbound

This paper cites RVT-2: Learning Precise Manipulation from Few Demonstrations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models RVT-2: Learning Precise Manipulation from Few Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.095191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.095191Z digest=sha256:eb2849f145e7059f09eaec1cb6caa026f83a44a49c71decaa916cbcf93ef4b61

Observation e1e07334-0fca-48ef-b91e-2331ec311f7d · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.187917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.187917Z digest=sha256:0b59998645ea3a290e00fdedcef8fde3b9fd07cbd67dbdf2e76a0dc9a86e122c

Observation b1fd0a23-d6f6-41c7-a9e3-d7f7e487a763 · outbound

This paper cites Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.268018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.268018Z digest=sha256:9ed519a948dd48936b9032de4537c5e02a6145cb3dc590477fc39a5cde7d0822

Observation 752060ec-438d-4d3c-ad9e-06c242fde276 · outbound

This paper cites StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.331513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.331513Z digest=sha256:7b1df67dfe44ba10ee8192920a7d5af697117e9b2cd1e95c8ab387ff0fa9083c

Observation c046748c-8091-4aee-bbdd-93269d49b7a9 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.983253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:19:48.393864Z digest=sha256:a5a890b3ecbf426b9eb86bcad05934a5fe7b34842e65ca1853bb2743b855aaa9

Observation cc182fc0-0a71-483b-a2dc-10046dc23220 · outbound

This paper cites QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.478639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.478639Z digest=sha256:627126abad5fb80b716de597686e09f2f48429f4d39e80236261ee072ffc109d

Observation 25f249e9-f40e-49e6-b0ef-f99092c85327 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.534851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.534851Z digest=sha256:4c88ddb97a9a2386a72be702f47239202ade33a815ae288387af140bad18ae6b

Observation 767aec74-c80d-4bca-8b52-692215445dc4 · outbound

This paper cites PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.590609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.590609Z digest=sha256:b4f39761a03af0191f1826efe1ce41b6139cda2b6fb9d7374f8d29c6bc8371a9

Observation a8669714-812a-40b7-8883-a0c45c288110 · outbound

This paper cites Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.667454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.667454Z digest=sha256:bb9289dbfd80e611bca771bb03059416ab54e42ef248251057bc7559f4f6dbce

Observation cd9d14f8-9612-487f-8ec8-2cde133f0f58 · outbound

This paper cites Peebles and S.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Peebles and S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.767277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.767277Z digest=sha256:4c6c82b355861308de31b67b874789142e7db649c60a3f1ab70ab679f3b4561d

Observation 06563ac3-47a8-4429-baa6-b30e995206c1 · outbound

This paper cites Radford, J.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Radford, J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.843593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.843593Z digest=sha256:c4e10f1b01ec40e4a6d79f11e4f1d465b6c6b2c98b5caf3764b7d96732135262

Observation d672b3e2-c805-4b58-81a8-5925afb5e535 · outbound

This paper cites Denoising Diffusion Implicit Models.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Denoising Diffusion Implicit Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:48.975765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:48.975765Z digest=sha256:255f76ac7fa57a9949fc32048d07468c36165ebf12e1d7df0670d2a45a92b364

Observation cc9f1646-a4c2-418a-b34c-61d6444f36c1 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.097273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.097273Z digest=sha256:cf637281a9bceabafbbac8bcb056ed9cb180c72c6e16d4e617aa7437a3f4dc57

Observation 23102af3-4a0e-486f-b583-ba540c05c04d · outbound

This paper cites Rombach, A.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Rombach, A

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:19:51.738404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:19:49.191361Z digest=sha256:8d7cfddec5e4c6ac0ea753371f5d027f5463202842dfacb2f9164c44a6aa98c7

Observation 1c149296-e834-440a-89e7-1cfb36079b83 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.316462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.316462Z digest=sha256:5be827115f216f4003c734350d569e5703c1a133edc23e7d23717722ed2e2f59

Observation 1a4a7068-291a-4154-a9ff-248e47bae85a · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.424569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.424569Z digest=sha256:bd6f6885b537c58c7e7bced0106f0de725e75ef4029b5dd15b123710d3def6b0

Observation 37bcff39-8ec9-46f8-838f-803038cf6f9b · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.501211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:19:49.557016Z digest=sha256:110109d1ee42319950c0111bb5caad4611cc26cfb9112292da35413f56bb22be

Observation d46dfc98-2a71-4997-be37-2ba63028c1af · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.656649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.656649Z digest=sha256:5374ba4af592c58f43419ea8cf8fa0cc949d5335be62c4bd9ec35f2bf3c4d40a

Observation ef3bd510-365b-4b64-a3ab-bce873a810f5 · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.758226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.758226Z digest=sha256:422f6522d38ebb5362b70684f39800c6b057824447f413de29a8f4735c799711

Observation c004a486-b24f-41c0-8c5f-aa26d8d09415 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:49.893880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:49.893880Z digest=sha256:4e65de24ec05dbc238deecf1352becd57f698d52411551ce1db31656d66ac06b

Observation 58340e5a-1563-48b3-8529-a965af4e892c · outbound

This paper cites an unresolved cited work.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:19:51.283184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:19:49.988923Z digest=sha256:f2b2e7c63f76393b9942e2be7540aef3aff07071d323d7692033254eb2f7646f

Observation 479ab36c-5a63-4a51-90f5-cdbe56196a57 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.092976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.092976Z digest=sha256:4468f2d0f1b5335d97618638028d999a077219cfd929ea07d3383c7eebf2aab1

Observation 6e930a4c-d7e0-4273-b5c3-18d9d8fce342 · outbound

This paper cites Kirillov, E.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models Kirillov, E

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.211537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.211537Z digest=sha256:d261dc69366bf3a8944272a708f18a8a93eca0be35ed34b70cc18f53360299b0

Observation c57b9e74-d5e7-4767-8b1f-30d0aa36d0b2 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:50.356389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:50.356389Z digest=sha256:8100c116d2117d103895fa93c1a0c3e5e77c473092d006c80edbc3a83c4f33ce

Pith citing papers

No inbound Pith citation observations are available.