Pith. sign in

Paper Citation Record · LEDGER

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

As of 20 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 24 inbound Pith citation observations for arXiv:2506.00123.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00123 v1

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:15.718768Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:47:29.605186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:28:34.757769Z

Reference resolution

100 of 112 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ee2d789-c94a-4f9b-a676-0cb0ed3ee942 · outbound

This paper cites GPT-4 Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.094257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.094257Z digest=sha256:e53ac2f5acb1f5a5b5d76fd09b71b8a8d12163f28247b49337a1e927ce1b5a7b

Observation 0674a851-2993-401b-9646-6fb3e9800e24 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.184567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.184567Z digest=sha256:71ff2c1ae1ebfc4272d5f7a5c5bad73e97e2d7550416cc95e38ce7f8e3821a01

Observation 8518c5f4-2e1c-4473-9f98-374b29f9bc36 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.271878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.271878Z digest=sha256:be09184fd967cd0d7039d3b02e8b839768aec46fbd99cb900e10f776a63a0456

Observation cb545fed-1e29-4216-b41d-865137d9a172 · outbound

This paper cites Claude 3.5 sonnet.https://www.anthropic.com/news/claude-3-5-sonnet, 2024.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Claude 3.5 sonnet.https://www.anthropic.com/news/claude-3-5-sonnet, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.351375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.351375Z digest=sha256:9b4b1a356a6f97106176b9f45fc8228b8c089dcb2d57878a82a44d1df9a0b802

Observation 0e05a99f-c97b-4d81-8e45-9c47448f95e1 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.414785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.414785Z digest=sha256:4cc5a7a7b28d2077115703ef9aff342510cebba9c741861dddf57690aa7c0250

Observation fc79de96-659e-4e20-85b5-93b6aabad814 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scanqa: 3d question answering for spatial scene understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.495975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.495975Z digest=sha256:ee72c305b0772dedce3ea50c746f136dc37bf7b4ef4627f359c4cfdd1c468e78

Observation e5e7458e-0d66-4212-a353-de4c5255c4d4 · outbound

This paper cites Qwen Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.566118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.566118Z digest=sha256:d21a01d8b9ff65c7aefee78e63cab8738a9b04a413cfcddb12b3a392434d28f3

Observation 3d9a7022-2f3a-407a-8a0b-fd1d9156406f · outbound

This paper cites Qwen2.5-VL Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Qwen2.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.626048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.626048Z digest=sha256:adafa51b20ad6179ecc79a6f08cf43a0f7985b833ea813cd0a2d2f0107b32793

Observation c32fc78d-0066-4baf-b491-4b4462ee1345 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.701912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.701912Z digest=sha256:e9587b7ebe42f01f21befdaf74ab4a06c0fe108e8a35b3ebcb7494c89bead33b

Observation 38f2732e-a4a2-4351-b414-6069c7abdd9b · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.777108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.777108Z digest=sha256:a7e09fd1d537cababf4b47384831994c6a5d4f369eec41fa79969b901380ae3e

Observation a0719468-b015-4571-b5ab-ce5ab0cee77c · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.844533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.844533Z digest=sha256:abe7868211ebb91799fd6f0291005c9766bdbdf38870659bcf58915db3a745ec

Observation 696b5cb8-bf54-4657-aabd-a0663a8265c3 · outbound

This paper cites Language models are few-shot learners.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language models are few-shot learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.930024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.930024Z digest=sha256:f3b46c1d3e13107d176f099c849c28a9b1f63717a9644bfcea9042d0cc75ef56

Observation 0b5be8d8-f62e-46fc-830e-a9989623dece · outbound

This paper cites Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.000071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.000071Z digest=sha256:3bc8103f81348452d261384ce51732e7fb41ddd5e1d3bea037cc3c3245d83161

Observation 105cfb30-dcb7-4e70-8984-d9e75ec5ad05 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.066166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.066166Z digest=sha256:fc7999fd9007a71f82adad23da5c3e908cfd550a538ad7967e47acaa300fc571

Observation 8675d496-772e-46a3-bca9-59a9e49115ab · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Eagle 2.5: Boosting long-context post-training for frontier vision-language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.119149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.119149Z digest=sha256:90b555e273674d31bf995c51585fb9448a84bebf7940a49ad43da5c0205c5e45

Observation 1e55fa0c-5216-4a00-976a-ae8b94f5868f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.213009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.213009Z digest=sha256:ba946846c80397f7fdf6beffc6183fab9c58cfead0f3b085a796a512e255c044

Observation a4876270-6c15-403f-a4a6-07c8cb24c66d · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Sharegpt4v: Improving large multi-modal models with better captions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.277793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.277793Z digest=sha256:8830ef615189cccdd8e14a0fb16f00a52e9d6a02e1c339f929d52753c90c15f2

Observation d1407922-9ad2-4d3f-a646-e4f4e64c2731 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.351955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.351955Z digest=sha256:b526e5ecc865450d487062346f30e57aed33335648282b6a2df29f4bc60ecc26

Observation 3984c22a-1ba4-4e61-8e53-46502b5649fb · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.427092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.427092Z digest=sha256:3eeaf12f7b34ceee969888da2f67136cd64b1bedfd0dbf797d5ac1838e823368

Observation faf511ce-2e13-437b-8378-3041049502fe · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.508644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.508644Z digest=sha256:29246d0c6cc124cc1fe7cc2667e24bf36d8eb629f90e18fa5e871b41e5548496

Observation 89df5c1a-f92d-4388-af89-4fa46b24b6bb · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.590815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.590815Z digest=sha256:200d76ed54885e79dea0d9583f22c2fb12fcc9cbf8a4f4372cb8d55753e9e22e

Observation 38816cc3-5d58-4c66-9782-0b54dcd0d5d0 · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.666040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.666040Z digest=sha256:05a9117f2f8bfae65c4f9d70771b59bb5f125d4010172f96c3de3956d321a0b0

Observation 75ef2d6e-407d-4877-85d3-9cfbbaa58328 · outbound

This paper cites Spatially-Aware Transformer for Embodied Agents.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Spatially-Aware Transformer for Embodied Agents

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:16:18.467967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:09.736172Z digest=sha256:9aab68c23df6397f8bd5e200b5b7e7983081f7514afd14be925fb5bcb8d786fd

Observation 80793765-2150-4de9-a05f-3f5f0f266ed7 · outbound

This paper cites Local all-pair correspondence for point tracking.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Local all-pair correspondence for point tracking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.802384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.802384Z digest=sha256:e63d6eaefb8dcf854239b78149a4c2a7b1a8d11658667cc56900c61103154ebc

Observation 8d538287-bf9c-432d-833a-3fc1eeb760c6 · outbound

This paper cites Simple and effective multi-paragraph reading comprehension.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Simple and effective multi-paragraph reading comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.857532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.857532Z digest=sha256:329f60e7aef4980033c009f24699ca99ef96b5ce76199aabe3f84ca6aa48431d

Observation d85376bb-fe9a-4889-b1b1-c7529397f013 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.945386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.945386Z digest=sha256:8a38e574c60c5975ba24c227191096e46d2eed2c92cc95b9ba6bedc19655f8bc

Observation d9cd709d-1052-43be-89e4-9d466b983348 · outbound

This paper cites Language modeling with gated convolutional networks.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language modeling with gated convolutional networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.021970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.021970Z digest=sha256:62865ca700a5dcaf033ec94af0d24e9d963ebdf04de7320e9b896aa386de48e0

Observation 2d053136-7bcb-43c0-81e8-7dbfb24ed116 · outbound

This paper cites Quar-vla: Vision-language-action model for quadruped robots.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Quar-vla: Vision-language-action model for quadruped robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.102883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.102883Z digest=sha256:b7ddb8a398b9cedf784187318cc9bb8ba0731364d844ea051c3b68aa43e92035

Observation e06e275e-ba11-47ba-aa0b-f6fc31c2dc26 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces An image is worth 16x16 words: Transformers for image recognition at scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.177298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.177298Z digest=sha256:a49cf2245c506d5621406b5d559b6c45f91f1036912efe13931e2dcebf251d11

Observation 6db3c8b1-01e1-4c50-9129-6fd6c48ccb2a · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces PaLM-E: An Embodied Multimodal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.243116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.243116Z digest=sha256:840e13381fb9c6c480f1054358ba344da4cfd9097d4d0efc2b768e47f4a7e08f

Observation 2a0ae299-5844-406f-b090-fe6419e2ce68 · outbound

This paper cites Centernet: Keypoint triplets for object detection.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Centernet: Keypoint triplets for object detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.287797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.287797Z digest=sha256:f1d6185fd8fea005789573ce117e1ffa7b18ce86d8209344a4d06825c2d99fd7

Observation f4629046-9ed0-43c9-80a3-4f6dfb2ca8a7 · outbound

This paper cites Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.334062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.334062Z digest=sha256:a36f51b663e061c535fc7b953b0ceba8e87159a5bf3d4d83858d267f989639ac

Observation a8224738-d488-4f4f-8a18-9997eb7fa620 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Eva: Exploring the limits of masked visual representation learning at scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.403215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.403215Z digest=sha256:c452ce2c2fe0d3226c557bce9651c4a78e0de9260fbdae9178d0027a01ae613a

Observation 7c64dda6-17a7-4488-b991-c933513eb539 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.499411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.499411Z digest=sha256:e657d0e6a1f4d358c06444ada598a17d534b9ce39c21230a8a259c7c3cbd5efb

Observation a21a4e9b-e31b-4fac-89a6-3bdabb2ecf9c · outbound

This paper cites Rlafford: End-to-end affordance learning for robotic manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Rlafford: End-to-end affordance learning for robotic manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.597240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.597240Z digest=sha256:4c30577ad53b8b4e46e513a3f2f49e958eeca28c335beb09457816b15d813d7e

Observation c49009cd-d63c-43e9-8140-b302563f7a46 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.648005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.648005Z digest=sha256:a88bc8f7d7e7fa178dca948fb6d3a7ab09493e58669b12c94ea9697aa3cfd74d

Observation 442877b5-7d93-4dc8-9674-df4b5cbf5740 · outbound

This paper cites Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.708199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.708199Z digest=sha256:37e3e88cecd3136bc126620e1e9aa902c33008a6292e797b2bb1bc2eb52fe4d3

Observation dee8501e-5fbc-4928-8892-5c0aed834f35 · outbound

This paper cites Girshick.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Girshick

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.791941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.791941Z digest=sha256:4285c11c0403f1d21b76a84b8788cf9541b825512cd8ea63297cc30de77be6eb

Observation d24edc9a-76f8-4ada-9938-76c194bc2301 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems , 36: 20482–20494, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems , 36: 20482–20494, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.876192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.876192Z digest=sha256:d99605421839889b76ccea216fd35aaa88e41a3fde4c1007e864832b29645872

Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.945513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.945513Z digest=sha256:5948a3abdcad8501df4a7cd145a8469eba8d3fcad4c26fe1ba6314855ac2e2d2

Observation 83ed7cf1-98a9-455c-92d6-50f5faa34813 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.043939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.043939Z digest=sha256:67465ea2c2db164cc91522442729e5f86ab58f80cc534e03e94024f7235b14df

Observation 10c4ad44-4a1e-4e28-a4ca-ca0c0ec860da · outbound

This paper cites Chat-3d v2: Bridging 3d scene and large language models with object identifiers.CoRR, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-3d v2: Bridging 3d scene and large language models with object identifiers.CoRR, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.134735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.134735Z digest=sha256:9ac1c2bf933d2493aa12938563e35e4814da302e846c17f816427c71775a0400

Observation cc41a651-4953-4a73-932e-78c486576128 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces An Embodied Generalist Agent in 3D World

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.217497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.217497Z digest=sha256:46ce1940b423034b0b0a564acb0e4a77116bf37d787a54eea44cfc320d445930

Observation c08a66b9-b8f2-4040-b0e6-1a299e52ec87 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.289544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.289544Z digest=sha256:890faa219f7e47839895bcd5f1436fa72e6a62424ab5ef6ae981be23d0c8f976

Observation 834ff20c-66be-4c9f-b9db-6b92c79dadcc · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.361561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.361561Z digest=sha256:4a1c26eb51b99f56cc579b644b6f2bf714c4b97a4c2a68f6d06e57c0002d9e08

Observation 77e4ec0b-04f0-4385-a4a5-3a983426f210 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.470944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.470944Z digest=sha256:8e7d67ebc4680d46788a20fa0dbf8769694703138ce91accd0b36936e5dcc178

Observation f7dafc7f-ff80-4db9-b147-1bce094aaacd · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scaling up visual and vision-language representation learning with noisy text supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.541136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.541136Z digest=sha256:03f661c5b73620364b166b2f26e4dbaba24a7582f75096ec39bb5dd48a0e202b

Observation 44b23ff9-3d8f-47d5-a5bf-9b174e55b26c · outbound

This paper cites A diagram is worth a dozen images.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces A diagram is worth a dozen images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.628090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.628090Z digest=sha256:08f345fbedab6f0c2aa5d9eac4fc1dda804520ddbe146ff7d8a4adc453f86f43

Observation bf0d064b-94e1-4314-b8cc-85003f6d3a3e · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OpenVLA: An Open-Source Vision-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.687388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.687388Z digest=sha256:c09fc782f01e48e387779d57391fffcf793e9c6cc7980099c5c014e2c122322c

Observation 22f767ef-56a2-4155-a22e-5bbff6497ade · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LISA: Reasoning Segmentation via Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.756541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.756541Z digest=sha256:060ce9c76fd68ece6fd05f80b8d5e745802bee39de588fb66bc2b8e8be8cfb34

Observation fe47318a-98c5-404b-89ef-ab3a454e34fc · outbound

This paper cites Cornernet: Detecting objects as paired keypoints.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Cornernet: Detecting objects as paired keypoints

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.826342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.826342Z digest=sha256:3d65d2e9372bc6ce4340c3ced4cc0a3d82962bb0128889ae7f14e6d801677668

Observation 89609035-90f0-40cf-a3e8-4f2913f97ebc · outbound

This paper cites Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436, 2018.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436, 2018

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.899475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.899475Z digest=sha256:9d6e13545e220aca5566da42b99b8a6ebb14b3fa7e6e7f80684c92d1e92f6fd7

Observation a87c0ae9-e60d-4481-af81-960fba18264f · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.985639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.985639Z digest=sha256:024ea419b1cfbce0b20a76b82f2189e1d2bb1a91872ec27a5a354a9cd881ef31

Observation fb31d28c-6296-4572-9949-15950c5d4bab · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaVA-OneVision: Easy Visual Task Transfer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.062017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.062017Z digest=sha256:95ad11691b5c61b69812c671b5b38fb0c5a4cdef2c5fdd692f37c948ba18ec5d

Observation 2d328658-9d72-4246-910b-1096cea908fd · outbound

This paper cites Learning agile skills via adversarial imitation of rough partial demonstrations.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning agile skills via adversarial imitation of rough partial demonstrations

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.141295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.141295Z digest=sha256:e3abb3107afc647ba2c1d3fab96a760c2be9dc55f7fc065abb7807322d7f9d4e

Observation 96670f96-a2ab-47ea-be94-64bf0c89c8af · outbound

This paper cites an unresolved cited work.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.205947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.205947Z digest=sha256:66b8bff3e911e290a34adfd61a84bab062e874cea2bc9652c1ce816fb8f9ced9

Observation e3e139c2-7507-4bfa-abf1-f3de348530d5 · outbound

This paper cites Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.285277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.285277Z digest=sha256:e6ec20e805cde6d17f3c6367e84660cc607ccac64660975e9054841ca0d91b8b

Observation b3524013-2008-4470-84d9-193b96830b34 · outbound

This paper cites Exploring Plain Vision Transformer Backbones for Object Detection.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Exploring Plain Vision Transformer Backbones for Object Detection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.363276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.363276Z digest=sha256:8ac3a3bea15e864faf2d8247db53e8bf7ec18d8ad03f1540ba4223307c1cacdd

Observation 2c551f63-1ff8-4ab4-af58-6f087d92ba59 · outbound

This paper cites Robotic Visual Instruction.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Robotic Visual Instruction

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.453068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.453068Z digest=sha256:125935b9856a44a202406e5567f89930a9e32dc85f7414774f8b6707a740e169

Observation 8f210c89-75f5-4c59-8221-675fa03e43a8 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Code as policies: Language model programs for embodied control

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.537280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.537280Z digest=sha256:cbe3d63256ae3cc7d66519ff23461bf2660272ac240ff69cdcc75bf571a4f93b

Observation 1f13fe24-c711-4692-bfd3-03c1f25cdbde · outbound

This paper cites Vila: On pre-training for visual language models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Vila: On pre-training for visual language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.611092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.611092Z digest=sha256:0f2f93b9535703b82ba24ccb20efb67c07513f91e7e5795107051c73c78682bd

Observation e8a3dc0b-fc1d-446d-ac2b-f7452a92d593 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.689514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.689514Z digest=sha256:2f129ac56cfad9b5c368e3e7fe4389bf8b48999ef842b471ceb546fae0ca7a83

Observation 06ca9977-97ca-4c6e-a229-b1cf9a2edc0a · outbound

This paper cites Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.740878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.740878Z digest=sha256:fd239d5a89e256dfa8edb2abab8e0bb73dda432585428090871efd75e02c4b34

Observation 3daa9635-42f0-475b-8fd9-4acffa328232 · outbound

This paper cites Visual instruction tuning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Visual instruction tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.799211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.799211Z digest=sha256:a96f51f7d0ca03784879866c7a7e1c4143ec86dff9623a39b13a582ec95eb7f7

Observation 36f2259d-718a-4b25-8a9e-45b2f0ae47fc · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MMBench: Is Your Multi-modal Model an All-around Player?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.881890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.881890Z digest=sha256:084bdec2a6f4a2962aa6e2d929591071e51cd2310d8592444362a2a7d99e770a

Observation 689a1580-0f37-41e1-a0c9-08676e08dfa2 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.944251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.944251Z digest=sha256:be0107ae187ef8ac3e96bd8c69c1e029e8c24e4a8be82f9006892ccaadeff0dc

Observation 1a306407-9e2e-46b0-b61d-48280076fd1e · outbound

This paper cites Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.878623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:13.046266Z digest=sha256:fd78d6825470841305b9fa635955b94e6624e5d6eed1c4e9e78ad226846be129

Observation 9e31acdc-ec99-411c-8518-ebe57bced002 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.117762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.117762Z digest=sha256:e27eb670a35b7466d4025985880f020ae8d74ebe7a8e2e811d8141675a8e00b5

Observation 8ef13d95-c7ca-4dad-b000-c89cc7399d84 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.209440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.209440Z digest=sha256:3cf704b6c0b01bb6efaa481691cb1d7275e720cf7f6c5923e7a75248f3812e4a

Observation 5e872cd2-3e60-48b5-b16e-e9c2b767fd75 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SQA3D: Situated Question Answering in 3D Scenes

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.281354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.281354Z digest=sha256:d9f83aa6c115c45d909eaf796c93e5a1251ec11496147e64f71bdde1f5beffcd

Observation 3213c8d1-cdf6-4d3c-ad23-d6aa78d04de1 · outbound

This paper cites Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems , 36:655–677, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems , 36:655–677, 2023

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.723933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:13.345167Z digest=sha256:2ae7a37aca981b00acd59ff77f4bb84cdcc4b5d687e566a1f41d4f7935a76a4d

Observation 0efa2d26-a987-441b-8041-4ffb0a5eeab3 · outbound

This paper cites Walk these ways: Tuning robot control for generalization with multiplicity of behavior.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Walk these ways: Tuning robot control for generalization with multiplicity of behavior

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.514359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:13.401055Z digest=sha256:a37f61d44c18cbf723171dd59ef89180849dd566ca4692e7f9c88b0ed08a2cc1

Observation 060f98ff-3b53-4e2c-a68b-e3d12ac5f5be · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.241222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:13.458663Z digest=sha256:8355e95385c1a4a01bb6fa0574b1f254b22d338c7df773f3ec81978c1d30ea4b

Observation 8129ca85-7683-47cb-b8ce-15aea4b9dbb6 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.059153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:13.513833Z digest=sha256:e954da403c51dd563d124b3093bd75bfd50ddb324685cc0a3c059097c7082e52

Observation 3251f2c2-46c5-4ac7-92ad-062f19267b9b · outbound

This paper cites Infograph- icvqa.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Infograph- icvqa

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.829698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:13.613661Z digest=sha256:68608f1055c019a6f8fb6aa5421e8c96426faeb844d9d34e57a8d1cd8ea8efec

Observation 453ef401-645e-4c86-b12e-a34b569110d1 · outbound

This paper cites QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.696485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.696485Z digest=sha256:0878784407de48986f5bbd605cf82ac8b7083234f002bb21ad211a815a14db28

Observation 93073953-f24b-452a-846b-123bb4f6078a · outbound

This paper cites Embodiedgpt: Vision-language pre-training via embodied chain of thought.Advances in Neural Information Processing Systems, 36:25081–25094, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Embodiedgpt: Vision-language pre-training via embodied chain of thought.Advances in Neural Information Processing Systems, 36:25081–25094, 2023

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.643942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:13.764571Z digest=sha256:d517acf5c3d503b283b53165e4b22e0f6e0d68a7993408cf7870a2bafaf1b159

Observation 20ec6a30-082c-4364-9530-38aabc8e2478 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces R3M: A Universal Visual Representation for Robot Manipulation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.840242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.840242Z digest=sha256:4ed470102136f979a07445686699109795c8f181db955e5448f0fec65b546a0e

Observation 59d906ba-aa13-49cb-bb2d-79eef539f927 · outbound

This paper cites Gpt-4o system card.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gpt-4o system card

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.324300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:13.906541Z digest=sha256:a3417784bf392ed2c92966e4ea1f3af1e9c5536fad6fc8afeb299f4275ff177d

Observation 6fa5edb8-db68-4971-9e85-c132ddab602c · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.992330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.992330Z digest=sha256:935c199eef5e17e5d132a7f6ae8b4a9a9805b7bcb57f7d9c4e6e673731a8b2f2

Observation 859eb22a-392a-4190-b6e6-45f594123c8b · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.104148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.104148Z digest=sha256:29df5306d8f71007a340a80c75e4e1cd956ffe5b0d106acfb8863817ae6acd90

Observation f0efc46c-7e1b-425c-8ae2-b32a44a583ca · outbound

This paper cites Improving language understanding by generative pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Improving language understanding by generative pre-training

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.114182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:14.201504Z digest=sha256:297cbe5ecf6eddb11fa427b6dc96d0fdc041daf15ade713043b65d8deba4624f

Observation 5c89e3d8-d4fe-4810-833e-1931d3ffc11c · outbound

This paper cites Language models are unsupervised multitask learners.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language models are unsupervised multitask learners

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.918889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:14.275120Z digest=sha256:8b7d49354fdd296eb7ad534f485e3b681184d73d7bb88a94a05e348a7278519b

Observation ee3335e7-0fd7-4dfd-b5a8-50e2e5a0eabb · outbound

This paper cites Learning transferable visual models from natural language supervision.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning transferable visual models from natural language supervision

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.665790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:14.329877Z digest=sha256:ed8f73ac3b8c1def3fa8cf453b501877f5fcc6c1c43fcfa9208f3fbd5bc12ccc

Observation cf7157ea-f21c-4296-8883-60c90dca772f · outbound

This paper cites Real-world robot learning with masked visual pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Real-world robot learning with masked visual pre-training

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.397881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:14.436516Z digest=sha256:f7f30c1f3eca6324782b6221ac0222e7ac3975dec266b44b747619b0543f99ca

Observation 9add7871-7f36-4f34-a621-f36582eb566d · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Cliport: What and where pathways for robotic manipulation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.194602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:14.584871Z digest=sha256:c6cfdf04a2916bb319dfb0340167dfb0cda8174aebc8f5756c9e3e9e592290d0

Observation 2a71a9fb-6724-476b-abb4-411f7b0e9b7c · outbound

This paper cites Towards VQA models that can read.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Towards VQA models that can read

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.971053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:14.692031Z digest=sha256:782d37a2ba4e71082c1bdf0aff1a9ba58bad1f63369e989bbb1cec6d7609c060

Observation e6059b5a-5b2a-4cda-8596-327f9ed638db · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.782958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.782958Z digest=sha256:6d9012bf0b7f7c0f509ef251456fcacf38dd7c8c821a6015625d8abb9ff6c77b

Observation 74674bf9-e778-4018-b380-b4b587df19cd · outbound

This paper cites SMART: Self-supervised Multi-task pretrAining with contRol Transformers.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SMART: Self-supervised Multi-task pretrAining with contRol Transformers

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.859943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.859943Z digest=sha256:345e4ded9703b44ecc3f5e5201f32f05ddf381d07dbd5afd6c9841b3a297e029

Observation 8ffe0a2d-f610-42fa-841c-d5d9ab45293f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini: A Family of Highly Capable Multimodal Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.937333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.937333Z digest=sha256:e468d5291b1039fd8f9a6914b11d6cfc3c03b1ac88f28ebecd815b16d3d6a071

Observation 34992c8b-5c2f-461a-9ebd-97bd33f16688 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.034425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.034425Z digest=sha256:843b8ee7ee64deb5d00d5d399578b141fc537ad56e9c4305775a34a5ec3686d2

Observation dfa936c2-1361-44c1-aca6-700a74758dfc · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini Robotics: Bringing AI into the Physical World

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.105152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.105152Z digest=sha256:e4e43389c0c0d223b1bbb1296cbe9f56d4ebe4ffd9a3ecf3af1571c19d9c20b1

Observation 61693193-d58e-4177-8a7d-ee10df47b7af · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaMA: Open and Efficient Foundation Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.176031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.176031Z digest=sha256:fbc6c2149780ad4a157c31811947d7153d8660ff8ff53d4b70308948d8be9762

Observation 06719993-6dd7-4eac-95f8-55b4e8a21ec3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.268308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.268308Z digest=sha256:b057026505dd205c690c2d3a38970473fe11064c0a4ff4972c0a3e735834e341

Observation a41ee9bc-cd25-455a-8f98-20c48cd04f0e · outbound

This paper cites The all-seeing project: Towards panoptic visual recognition and understanding of the open world.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces The all-seeing project: Towards panoptic visual recognition and understanding of the open world

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.827901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:15.352827Z digest=sha256:58a1a63df076068b66b0c2136a246dc1ea15a00c8ac18ac4fa8c501c6a43d704

Observation e1d03a9a-3438-4690-a6a9-9da73d0da9e6 · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.449497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.449497Z digest=sha256:47ad00eee0fe038035c6daa75ead2e3a568021721b5d64e5215fb119744b1954

Observation e412d259-4bc0-4a4d-8e91-d339d6ca3b4e · outbound

This paper cites Grok-1.5 vision preview.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Grok-1.5 vision preview

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.721585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:16:15.510411Z digest=sha256:18c87c0be311114ec9b437e9b7712f1056e627b24c6bd309ad09543a18981f11

Observation eaff1423-9048-4f1f-8ee2-dd20b974265f · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.576937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.576937Z digest=sha256:b672596d00e695059e1cb1b6c658f03734f36b199830da9b66b509201884f54e

Observation 239354fb-65f5-4647-9534-f23f80f943f4 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.641523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.641523Z digest=sha256:efd1db60f47f03a5a3b2c068b57e92326d9957264d9198d14fe1a25a20993255

Observation 8944cc7b-10b1-4c98-b1bf-5159dde0e1f4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.718768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.718768Z digest=sha256:f7dfd5ad0c17f47c030870ca6b29a2b0af2e9c15f043c0684e1d034f851c4b43

Pith citing papers

Observation 92387108-7980-4289-8454-b5ce4b3b9582 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.605186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.605186Z digest=sha256:a5b9dee621e7a906d60d8bf650cd05ee62bd02a581f49277a288977493c61f05

Observation 6253ce52-1ae4-49db-9f13-5e9949c949b3 · inbound

Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs cites this paper.

Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:29.157371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T07:37:38.901354Z digest=sha256:3fabec14213cadd8ff22e414d5a8c0c3f4162351d7f93b1bff64c89abd00afbe

Observation cc80ffac-da6b-477b-bd31-f75cdccac3fc · inbound

The high-speed X-ray camera on AXIS: design and performance updates cites this paper.

The high-speed X-ray camera on AXIS: design and performance updates Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T18:48:03.418621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:48:03.418621Z digest=sha256:857ec95e71bf397629b7d7575af3a8d5e3ba4ef43aa22bb7b0f736fc17f082a8

Observation 8684ed19-c732-44f0-b64a-6561ec1fe91b · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.932813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.932813Z digest=sha256:33ab0acc3401d0847427f314f06c975f5dd0e1e778be8062bfbb79c547fd5c1c

Observation 80a4f563-9ab0-4a4d-96c2-d8e8d525ba01 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:39.833644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:5dc7a8b040dfbdc829f3d98c81fdaae080a10ac5a63c4736ba892bd012664b94

Observation 35c55ac7-b8a4-4f5b-ab32-c4361e29ce64 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.714505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:c8283f81f7c1cb9116de9c09f469b98868806ecd090f1f6191df56d68c0df02d

Observation 189a4e30-35d9-47ed-9375-7da26e19f9ba · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.373521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:8578d848948901a251ff3b3812facbee87436161d7ecd0846e4ee3fbca7588ce

Observation a7ea5a71-deb2-4e97-9c2b-897a4aad0f53 · inbound

3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding cites this paper.

3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:56:00.254914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:51:52.063238Z digest=sha256:1c07af3a2cf776b1619db42306777ca9bacb3c3a89ac2c41254f17ff2303856d

Observation 0749cb87-7f03-44c0-bac1-e3af4955cb14 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:58:49.605881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T17:57:00.909897Z digest=sha256:878ffcbb758dc0de294685083601365c8597e61c44252a4f1757e836d1adc6a1

Observation 37aa5277-49f8-4b04-8a43-0d65279cff87 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.582050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T19:02:05.125937Z digest=sha256:aa6784a7de8c7bdb653befa3d38b8188c83994ff0c8a2c454d55bfe2d2ae8da0

Observation 319808e9-17f6-4ce6-90d6-8db16d291176 · inbound

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation cites this paper.

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:24:40.387049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T06:24:30.669956Z digest=sha256:35b6cd630b918ea5adadbb94ca411de5be2bc38c8b5ed7027371001ea39bd08b

Observation 63e8f605-6c9f-4107-bda1-eca7581de84b · inbound

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation cites this paper.

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.282353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T17:07:49.005419Z digest=sha256:cd74da54b78da166ce29807fa1e64a1c3c7694171a0d7841a2d921c3df666c79

Observation f8731dac-224d-47d0-9833-45ab6014a0b7 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:58.950221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:67bf83c773497cc193be15d8bbd09f308b748dc42e4b53d1f093857045e3886b

Observation b4b13735-5d8f-4332-935e-4cce4db729cf · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.937015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:820dc66aac0ab074ad939e65a843395e54fe0bbbb21d878f2d3d2e261a3f0b14

Observation cdae8602-7c06-4f94-b73c-6ab9c10f8ddd · inbound

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs cites this paper.

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.075845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T11:02:07.122615Z digest=sha256:fe1a22fd9f9ca87080500453eca13307ddab76a2b295119ab3f7bbb22d9a825f

Observation 8b86d7ef-1904-4510-9c25-e011c7353b06 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.154508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:8d412024b455277977114bb1bd19664e5fe84b0a60f85fc04e8d2f4797622934

Observation ad9c88a9-defe-4f09-9434-36ed67cab158 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:4332c36c2bd1733285d7e9f14085f808a700725675f9da0418cfbb53bb2c24b1

Observation be10cee9-94d9-40db-a399-91baf0cabd54 · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.760148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:26:32.209719Z digest=sha256:d55251c446d2a0b6d2ea5e901c3a9980a7e2701c9a70c1b9ec12a5801f32f9cb

Observation ca74376b-ee0f-4340-892e-ef46ed4806a6 · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T11:43:49.639057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:43:49.639057Z digest=sha256:7fd350fa006c186f5de8cbcc2416fd83d90020879ee9da444eb8f79b25b77d41

Observation e7ce9633-15b2-48d6-adad-d5b498358984 · inbound

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale cites this paper.

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.351453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:28:14.694168Z digest=sha256:bf883ba9c265cd91154d145210a7d78c43570a9b0934bf29aaef6162c14477f0

Observation 880ac1a5-6849-4c80-8dfe-753e815bd6bb · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.366106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T11:17:48.279808Z digest=sha256:490ae69b6e891a5413bc6076c22c13b37530ea437878a37c09854aacdc06b209

Observation 9d610a94-38fe-468c-b91c-69837c2cc378 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T11:20:26.553333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:20:26.553333Z digest=sha256:e4ed7cbb78f0fb6c1405126f5b3c17bbbed41c70848b4d3bdbcdfaacf6e40488

Observation 5ca3a554-d397-484d-9748-597a412c8479 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:1480d62a1d74a5f5e38d8e2be0f28cd829a036e6d539bc3bc558d98cf0c8997d

Observation e0899ac6-ff62-4be6-bdbd-271e62038c2b · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:39.709300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:39.709300Z digest=sha256:7edcc53677db53835c15aac37499af64041dc5246138980c2a889243fa28ac1d