Pith. sign in

Paper Citation Record · LEDGER

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

As of 7 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 24 inbound Pith citation observations for arXiv:2506.00123.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00123 v1

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:15.718768Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:47:29.605186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:28:34.757769Z

Reference resolution

100 of 112 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ee2d789-c94a-4f9b-a676-0cb0ed3ee942 · outbound

This paper cites GPT-4 Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.094257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.094257Z digest=sha256:5b51dc7c46ea09aa8bbeffe532b044f1348b95275632589e5bea06d19802e10c

Observation 0674a851-2993-401b-9646-6fb3e9800e24 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.184567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.184567Z digest=sha256:330d94ccff10f91bc693acb9ea39af7f43bf6d7f0ff960ffbc0b0c6dd8dd5c05

Observation 8518c5f4-2e1c-4473-9f98-374b29f9bc36 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.271878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.271878Z digest=sha256:2b1a280b474e91ae6291c7e935e35839cb4e3489d55854482eb6a43b726a7a91

Observation cb545fed-1e29-4216-b41d-865137d9a172 · outbound

This paper cites Claude 3.5 sonnet.https://www.anthropic.com/news/claude-3-5-sonnet, 2024.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Claude 3.5 sonnet.https://www.anthropic.com/news/claude-3-5-sonnet, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.351375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.351375Z digest=sha256:8046f6a467510030ea0f229a7f2f2bdddd6eee91b8846cdbf53586c89e55d679

Observation 0e05a99f-c97b-4d81-8e45-9c47448f95e1 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.414785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.414785Z digest=sha256:31b4cf61ac0e497204e5c123fa0a6bf5eab9ef824dead51f1685ad364717ffb0

Observation fc79de96-659e-4e20-85b5-93b6aabad814 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scanqa: 3d question answering for spatial scene understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.495975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.495975Z digest=sha256:1e0c302db6353af016baf5d62c07accc260d1050ce0c2cd1bad10416932b27ee

Observation e5e7458e-0d66-4212-a353-de4c5255c4d4 · outbound

This paper cites Qwen Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.566118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.566118Z digest=sha256:88e558c980271f339587be867f6172e7fb4384bd43da4f76553e21a47467b701

Observation 3d9a7022-2f3a-407a-8a0b-fd1d9156406f · outbound

This paper cites Qwen2.5-VL Technical Report.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Qwen2.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.626048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.626048Z digest=sha256:d43c0d0e37110f16377cf902cbc83be8d26a8a9392770f4fee61fb874a714472

Observation c32fc78d-0066-4baf-b491-4b4462ee1345 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.701912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.701912Z digest=sha256:ecf95f6ae3ec20200d0329d4270cf4cd0a5f90e0c4141ceec797c8a493eef14e

Observation 38f2732e-a4a2-4351-b414-6069c7abdd9b · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.777108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.777108Z digest=sha256:b7efe03bfd731f83d3b58485c4fbc252df9a977fb2a128e10e58ff87be0e58c8

Observation a0719468-b015-4571-b5ab-ce5ab0cee77c · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.844533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.844533Z digest=sha256:943bf4b8a9742d68943c0367e33f7ea223092a97685b9d3c631602a77292e313

Observation 696b5cb8-bf54-4657-aabd-a0663a8265c3 · outbound

This paper cites Language models are few-shot learners.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language models are few-shot learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:08.930024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:08.930024Z digest=sha256:6b1e398d2d9fc00e9313cd5c4d61036c0fac31db2f31f0d45620d1cfe2006bfb

Observation 0b5be8d8-f62e-46fc-830e-a9989623dece · outbound

This paper cites Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.000071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.000071Z digest=sha256:1ad002fc8631e84c1eb88542ffe177d34b5ede926c9093d07dd1224c826a4dcd

Observation 105cfb30-dcb7-4e70-8984-d9e75ec5ad05 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.066166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.066166Z digest=sha256:add9a511e5935934c46b218ac66d639a6a1e1002982aefbb1bc275eae0a83f4f

Observation 8675d496-772e-46a3-bca9-59a9e49115ab · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Eagle 2.5: Boosting long-context post-training for frontier vision-language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.119149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.119149Z digest=sha256:664834415387b91f8ecfb94f021b6d927af011c7a651dc803402b4c48dfb505b

Observation 1e55fa0c-5216-4a00-976a-ae8b94f5868f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.213009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.213009Z digest=sha256:8a6bc502e2f4907b87365ebee77f29ba58362cd5a42564f0be2babb32fb94a81

Observation a4876270-6c15-403f-a4a6-07c8cb24c66d · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Sharegpt4v: Improving large multi-modal models with better captions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.277793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.277793Z digest=sha256:c0e2c57faad5e87c6855f9f68ac39700e144e739c3901e87a061521747992dbb

Observation d1407922-9ad2-4d3f-a646-e4f4e64c2731 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.351955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.351955Z digest=sha256:f6fbf6473fce4b0035450a84f5ccde9331fa0184feaeab5ab73556f64a0b312a

Observation 3984c22a-1ba4-4e61-8e53-46502b5649fb · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.427092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.427092Z digest=sha256:ff0851748cc6d5350dca47f9a321d4dd231c43a6564d8b24415ffd5d57a3ef8f

Observation faf511ce-2e13-437b-8378-3041049502fe · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.508644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.508644Z digest=sha256:3cf1f36396df4963cc6f758c79cd43690ca01f55754ad610a02c2b91115722ff

Observation 89df5c1a-f92d-4388-af89-4fa46b24b6bb · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.590815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.590815Z digest=sha256:e8ade7786bebfebd07209d679e86f69a373c4dacc062f1361eba657336d2f752

Observation 38816cc3-5d58-4c66-9782-0b54dcd0d5d0 · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.666040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.666040Z digest=sha256:e002af4f287d055cb1d489f32087c67820114d96815284892cbdb19a488f2a06

Observation 75ef2d6e-407d-4877-85d3-9cfbbaa58328 · outbound

This paper cites Spatially-Aware Transformer for Embodied Agents.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Spatially-Aware Transformer for Embodied Agents

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:16:18.467967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:09.736172Z digest=sha256:ca529b42846aeee9c50bf2dfc05a934ef1ba72df0a3638295ba36dc59612dfb6

Observation 80793765-2150-4de9-a05f-3f5f0f266ed7 · outbound

This paper cites Local all-pair correspondence for point tracking.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Local all-pair correspondence for point tracking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.802384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.802384Z digest=sha256:4a5118de67cc1a09949b5c739ee81c64ee309c8add7e8a9dbafb16030c27a1f8

Observation 8d538287-bf9c-432d-833a-3fc1eeb760c6 · outbound

This paper cites Simple and effective multi-paragraph reading comprehension.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Simple and effective multi-paragraph reading comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.857532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.857532Z digest=sha256:b68d99e717fd368b4fbb1ab910242fed95a12f9dd0786977dafc7f2b2ad78422

Observation d85376bb-fe9a-4889-b1b1-c7529397f013 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.945386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.945386Z digest=sha256:c1fe6ca701bc7bbef0c73ff4f817cfdce68da7b2e1a21b86f5ee988246981e6f

Observation d9cd709d-1052-43be-89e4-9d466b983348 · outbound

This paper cites Language modeling with gated convolutional networks.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language modeling with gated convolutional networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.021970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.021970Z digest=sha256:b3749890d4c09cc3aaa552f7dd0e4313b51af5f356728bd2d20f3c5c8a3cf4de

Observation 2d053136-7bcb-43c0-81e8-7dbfb24ed116 · outbound

This paper cites Quar-vla: Vision-language-action model for quadruped robots.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Quar-vla: Vision-language-action model for quadruped robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.102883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.102883Z digest=sha256:a4a3769887eab001d26196d139236827995085796376a22c86b9f44140e97e02

Observation e06e275e-ba11-47ba-aa0b-f6fc31c2dc26 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces An image is worth 16x16 words: Transformers for image recognition at scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.177298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.177298Z digest=sha256:9ef4d2347e105644b06af3b3ddea1cdcfd256e0cea77cdf18da2d2e82071ec19

Observation 6db3c8b1-01e1-4c50-9129-6fd6c48ccb2a · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces PaLM-E: An Embodied Multimodal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.243116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.243116Z digest=sha256:d959fa3fe3a1d53d03f4836b4cd7e0e5c26f6f62609fa6b623f8fc86e0b49b95

Observation 2a0ae299-5844-406f-b090-fe6419e2ce68 · outbound

This paper cites Centernet: Keypoint triplets for object detection.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Centernet: Keypoint triplets for object detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.287797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.287797Z digest=sha256:78a9ace5b5f574cba17fc9449f2b6489513e51784b7367522b851617ff788fb5

Observation f4629046-9ed0-43c9-80a3-4f6dfb2ca8a7 · outbound

This paper cites Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.334062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.334062Z digest=sha256:def242fda3bac288f272eb07eaaf342416de1c057a5804b743e3705b256e2d8e

Observation a8224738-d488-4f4f-8a18-9997eb7fa620 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Eva: Exploring the limits of masked visual representation learning at scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.403215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.403215Z digest=sha256:0fce50f807dc1fc479afc6494768e87fab0b8f60515d50b911dc91877030356b

Observation 7c64dda6-17a7-4488-b991-c933513eb539 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.499411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.499411Z digest=sha256:5fa5a02ee6ad877b5556e6812f9e4d09ecdc3bc98f8189036a279a6c76ba807c

Observation a21a4e9b-e31b-4fac-89a6-3bdabb2ecf9c · outbound

This paper cites Rlafford: End-to-end affordance learning for robotic manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Rlafford: End-to-end affordance learning for robotic manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.597240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.597240Z digest=sha256:bb2ea19038212655057b5278e7e46842ce0f5e525f0031990dc497b6d1e33e0f

Observation c49009cd-d63c-43e9-8140-b302563f7a46 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.648005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.648005Z digest=sha256:efc20fb6fdf037083f353ba0e2a17a835e37ca8cd6d5e3d07d08d6b83338a898

Observation 442877b5-7d93-4dc8-9674-df4b5cbf5740 · outbound

This paper cites Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.708199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.708199Z digest=sha256:bc85e3da4e1f0e8b5874891d453e562c898d8283a2b05ba9f2574b5e9e5c218b

Observation dee8501e-5fbc-4928-8892-5c0aed834f35 · outbound

This paper cites Girshick.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Girshick

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.791941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.791941Z digest=sha256:712d039cd0ba0f7f82aa260dbe03218c7d676fabe72447b623f900cb7919e297

Observation d24edc9a-76f8-4ada-9938-76c194bc2301 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems , 36: 20482–20494, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems , 36: 20482–20494, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.876192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.876192Z digest=sha256:97e172d0efda2e8e05cd7fb72f6eb007c6e27821f4f7592c07ba75029e9a7150

Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.945513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.945513Z digest=sha256:bf172191dc02fd10fd117e45e583ae3232dfc3c1d17448b2b03ee32d1acb14d4

Observation 83ed7cf1-98a9-455c-92d6-50f5faa34813 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.043939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.043939Z digest=sha256:b96c5ac0ea1e4f3747850630d46f0bc4fbab14fd5769465954a8f368bf010295

Observation 10c4ad44-4a1e-4e28-a4ca-ca0c0ec860da · outbound

This paper cites Chat-3d v2: Bridging 3d scene and large language models with object identifiers.CoRR, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-3d v2: Bridging 3d scene and large language models with object identifiers.CoRR, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.134735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.134735Z digest=sha256:d767b4e46ba57058fc42bdb64d11b970a0fdd556bb1dbce4cfe5125aeaf0d888

Observation cc41a651-4953-4a73-932e-78c486576128 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces An Embodied Generalist Agent in 3D World

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.217497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.217497Z digest=sha256:2f80c238fdd485c573863723669d204cf762c8d7e04597bcd9af2392cb777013

Observation c08a66b9-b8f2-4040-b0e6-1a299e52ec87 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.289544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.289544Z digest=sha256:91225fb1b7be51b869dda48d51358ca0655e6170b6a3abd46b0a46d05ec398d0

Observation 834ff20c-66be-4c9f-b9db-6b92c79dadcc · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.361561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.361561Z digest=sha256:a51db76b4216313eb9a906ddfc19dea2cea88119ed379c65db43b98701e7e304

Observation 77e4ec0b-04f0-4385-a4a5-3a983426f210 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.470944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.470944Z digest=sha256:75a3721db083923e4a9bcb08c7856fd499b8080ae427fbbf2ef3f1cbcc77c595

Observation f7dafc7f-ff80-4db9-b147-1bce094aaacd · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scaling up visual and vision-language representation learning with noisy text supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.541136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.541136Z digest=sha256:e1f2001392c9bcdbeaee5448dcc524d9029384ed633986b7b4a92fc1c8c068e0

Observation 44b23ff9-3d8f-47d5-a5bf-9b174e55b26c · outbound

This paper cites A diagram is worth a dozen images.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces A diagram is worth a dozen images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.628090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.628090Z digest=sha256:c8b2a142bdc8e14b6426412bab9044933ce29ca063bc311748b456162bc075a4

Observation bf0d064b-94e1-4314-b8cc-85003f6d3a3e · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OpenVLA: An Open-Source Vision-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.687388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.687388Z digest=sha256:eda24bdb9df9ba5645ee9e5533bfe619272da603e47e2f6e204dbd158f8bdf37

Observation 22f767ef-56a2-4155-a22e-5bbff6497ade · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LISA: Reasoning Segmentation via Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.756541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.756541Z digest=sha256:663c0a1ae58285ae3d4a98eb8683926d970ce45324a067bd5ef929e4447e5cd5

Observation fe47318a-98c5-404b-89ef-ab3a454e34fc · outbound

This paper cites Cornernet: Detecting objects as paired keypoints.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Cornernet: Detecting objects as paired keypoints

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.826342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.826342Z digest=sha256:ade2c245c8342e9acf4b19b2575badedb5d0268ea62da68bba0658758fddd998

Observation 89609035-90f0-40cf-a3e8-4f2913f97ebc · outbound

This paper cites Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436, 2018.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436, 2018

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.899475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.899475Z digest=sha256:f2a99524b20f5999225e306678257c1beeb2f65d73b3b59b44860887f9a671d3

Observation a87c0ae9-e60d-4481-af81-960fba18264f · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.985639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.985639Z digest=sha256:3b174aa0dc18ef7a34941568febd1ad7c57a142722901f92ed700be0752913c6

Observation fb31d28c-6296-4572-9949-15950c5d4bab · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaVA-OneVision: Easy Visual Task Transfer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.062017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.062017Z digest=sha256:ebb1be1e9b38544cc9cb1bbb289bfe9982cdb9fc3f8710db79ff92144c8a6c94

Observation 2d328658-9d72-4246-910b-1096cea908fd · outbound

This paper cites Learning agile skills via adversarial imitation of rough partial demonstrations.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning agile skills via adversarial imitation of rough partial demonstrations

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.141295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.141295Z digest=sha256:8f0e114fe9b9d2b1bc3d82ec6efa72216f50cf668abe8eca7f237d9a3ee235a2

Observation 96670f96-a2ab-47ea-be94-64bf0c89c8af · outbound

This paper cites an unresolved cited work.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.205947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.205947Z digest=sha256:cdcdac9386056fb61efb9095b3bbe4294fdcf03ca197673f4a56c6a3a1953e04

Observation e3e139c2-7507-4bfa-abf1-f3de348530d5 · outbound

This paper cites Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.285277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.285277Z digest=sha256:0751045cec380458aa9338f4e93cd0e3fd5216abedab95e951396fb53c2371d7

Observation b3524013-2008-4470-84d9-193b96830b34 · outbound

This paper cites Exploring Plain Vision Transformer Backbones for Object Detection.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Exploring Plain Vision Transformer Backbones for Object Detection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.363276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.363276Z digest=sha256:1f5777dc96cfac3a2fc47b1d701056c436f9c0aa3d66033fd828e8f8a2b2b045

Observation 2c551f63-1ff8-4ab4-af58-6f087d92ba59 · outbound

This paper cites Robotic Visual Instruction.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Robotic Visual Instruction

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.453068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.453068Z digest=sha256:a2151ef5160b742d7f348dc3c47a8d67a15a4fd4e49c857dfee6b639b2ee0017

Observation 8f210c89-75f5-4c59-8221-675fa03e43a8 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Code as policies: Language model programs for embodied control

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.537280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.537280Z digest=sha256:0424e3caf27510f5b5f5933e3e4003c997755f8251274a22b38d61df2eb58ff2

Observation 1f13fe24-c711-4692-bfd3-03c1f25cdbde · outbound

This paper cites Vila: On pre-training for visual language models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Vila: On pre-training for visual language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.611092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.611092Z digest=sha256:0b6766ebd319c0bf1e02ad9a4548abc95630a485c1f8c6370e6a8a9822345a03

Observation e8a3dc0b-fc1d-446d-ac2b-f7452a92d593 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.689514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.689514Z digest=sha256:fbbe2aa90e148614fe99cad3d41db3640a058518d62253386700ef68358f0adb

Observation 06ca9977-97ca-4c6e-a229-b1cf9a2edc0a · outbound

This paper cites Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.740878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.740878Z digest=sha256:1dee6506bdfcc15a3524a1c235a42326ea1ff308e54491ffd058278bf70e6af4

Observation 3daa9635-42f0-475b-8fd9-4acffa328232 · outbound

This paper cites Visual instruction tuning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Visual instruction tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.799211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.799211Z digest=sha256:88a4ffd7c85b46a514b71e919366a173a0d6cc3faf4838875855c97a2114dd60

Observation 36f2259d-718a-4b25-8a9e-45b2f0ae47fc · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MMBench: Is Your Multi-modal Model an All-around Player?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.881890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.881890Z digest=sha256:de20ac643813e15108b84ae71e83ec6f2f106e0e8474a9628d0f1e19c86a17e6

Observation 689a1580-0f37-41e1-a0c9-08676e08dfa2 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.944251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.944251Z digest=sha256:93f4a7697035994bd8c0f796d74102be5473091f4ec674c1a9b98538875f2d08

Observation 1a306407-9e2e-46b0-b61d-48280076fd1e · outbound

This paper cites Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.878623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:13.046266Z digest=sha256:298c0df9bf93d5c952277ad6c1e6b21b53e595810f5e286de67368cb08b5d5e1

Observation 9e31acdc-ec99-411c-8518-ebe57bced002 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.117762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.117762Z digest=sha256:aa3279292115820662818eb5396f594630f4a97c1d49c28d7cce0b50ac9928f3

Observation 8ef13d95-c7ca-4dad-b000-c89cc7399d84 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.209440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.209440Z digest=sha256:076d7c4d1b5c7e6395b504eedf7125c72de523bef9732fada844bc049fdae968

Observation 5e872cd2-3e60-48b5-b16e-e9c2b767fd75 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SQA3D: Situated Question Answering in 3D Scenes

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.281354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.281354Z digest=sha256:19479e0a700e7d51061428fcbf6e89e6cb681219cc4a395978f3616ae80e8d1f

Observation 3213c8d1-cdf6-4d3c-ad23-d6aa78d04de1 · outbound

This paper cites Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems , 36:655–677, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems , 36:655–677, 2023

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.723933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:13.345167Z digest=sha256:24bce9ffafeef1aa93a13b380b586c8339347b6d2d2fd9e8aa5bd3ae3bdf0a71

Observation 0efa2d26-a987-441b-8041-4ffb0a5eeab3 · outbound

This paper cites Walk these ways: Tuning robot control for generalization with multiplicity of behavior.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Walk these ways: Tuning robot control for generalization with multiplicity of behavior

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.514359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:13.401055Z digest=sha256:9a756d3dca4fc4d039fa1b10084f25a987fa5a8f3c920e16b40808b90ca62cb0

Observation 060f98ff-3b53-4e2c-a68b-e3d12ac5f5be · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.241222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:13.458663Z digest=sha256:3fce3542a8191612a59e37e1ef79c04e8e143ee17f777eedcd0f7ce06c1cb78a

Observation 8129ca85-7683-47cb-b8ce-15aea4b9dbb6 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:22.059153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:13.513833Z digest=sha256:2dadddb47be9e4c2c30939843f15b1627c86cbcde140e44886451216badb6fc4

Observation 3251f2c2-46c5-4ac7-92ad-062f19267b9b · outbound

This paper cites Infograph- icvqa.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Infograph- icvqa

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.829698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:13.613661Z digest=sha256:f50f53856c716f018072c78f709bbca5557c5213b8fb23201e7010fc3f3c907e

Observation 453ef401-645e-4c86-b12e-a34b569110d1 · outbound

This paper cites QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.696485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.696485Z digest=sha256:598ff8af177e4f07720cd29732606603fd118e3e23b762c6dc81ba302e13988b

Observation 93073953-f24b-452a-846b-123bb4f6078a · outbound

This paper cites Embodiedgpt: Vision-language pre-training via embodied chain of thought.Advances in Neural Information Processing Systems, 36:25081–25094, 2023.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Embodiedgpt: Vision-language pre-training via embodied chain of thought.Advances in Neural Information Processing Systems, 36:25081–25094, 2023

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.643942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:13.764571Z digest=sha256:fcb468a5b1a069cc7ae1d6916036158d2b51eeab8d401fdca4b154f8082961d3

Observation 20ec6a30-082c-4364-9530-38aabc8e2478 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces R3M: A Universal Visual Representation for Robot Manipulation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.840242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.840242Z digest=sha256:deb1d4a4a5ad0ef6edce026b27a737cf802254a7b39ae91e1c9fbcd429977622

Observation 59d906ba-aa13-49cb-bb2d-79eef539f927 · outbound

This paper cites Gpt-4o system card.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gpt-4o system card

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.324300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:13.906541Z digest=sha256:c2bc9107c22dc26fe275160d30d5290820c521b98d4ffcb28b25118d96f5a3b1

Observation 6fa5edb8-db68-4971-9e85-c132ddab602c · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.992330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.992330Z digest=sha256:8deaf882d320d0bbb3d3ee8dce6ee64270ed495ebc35baf0cf6b578f2eb18a65

Observation 859eb22a-392a-4190-b6e6-45f594123c8b · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.104148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.104148Z digest=sha256:e4cec2d1d2f1c91cf1bedb1ea2da857eb871cdb88fc2f954e8f8ba32061c45fb

Observation f0efc46c-7e1b-425c-8ae2-b32a44a583ca · outbound

This paper cites Improving language understanding by generative pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Improving language understanding by generative pre-training

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:21.114182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:14.201504Z digest=sha256:17b83a9cacaf8b7acf121c3427826ef032e8bb26aa60efbe9b18f1d4a6cc31c1

Observation 5c89e3d8-d4fe-4810-833e-1931d3ffc11c · outbound

This paper cites Language models are unsupervised multitask learners.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language models are unsupervised multitask learners

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.918889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:14.275120Z digest=sha256:71255034cc908f4771de145a60606cd0ccbfaaff3877bae025e3590804ae5135

Observation ee3335e7-0fd7-4dfd-b5a8-50e2e5a0eabb · outbound

This paper cites Learning transferable visual models from natural language supervision.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning transferable visual models from natural language supervision

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.665790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:14.329877Z digest=sha256:e8db22dc70866ce52babf8bb29d9f2e084b37652d208f201870fe438b1ab474c

Observation cf7157ea-f21c-4296-8883-60c90dca772f · outbound

This paper cites Real-world robot learning with masked visual pre-training.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Real-world robot learning with masked visual pre-training

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.397881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:14.436516Z digest=sha256:c25c528fdf17e21c68c4ab81ee9fd1419478e926689c9c4d27feed120db1a97e

Observation 9add7871-7f36-4f34-a621-f36582eb566d · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Cliport: What and where pathways for robotic manipulation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:20.194602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:14.584871Z digest=sha256:107ff1bf31579c716b1be0ae9c33d97aca5a5531d013c53137ab5aa6399f0ac7

Observation 2a71a9fb-6724-476b-abb4-411f7b0e9b7c · outbound

This paper cites Towards VQA models that can read.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Towards VQA models that can read

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.971053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:14.692031Z digest=sha256:833ca8123aa3a582b7873676fff89371a7e259dfbea58004182cc668e80e87dc

Observation e6059b5a-5b2a-4cda-8596-327f9ed638db · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.782958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.782958Z digest=sha256:5f8740f8467c2a39792b02257cc5a11f043a8618d55920a8b0e07ba2685b3a10

Observation 74674bf9-e778-4018-b380-b4b587df19cd · outbound

This paper cites SMART: Self-supervised Multi-task pretrAining with contRol Transformers.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SMART: Self-supervised Multi-task pretrAining with contRol Transformers

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.859943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.859943Z digest=sha256:0572d0da6d87143ab05838d1c15c5eb92b52617114c568a788cc76db16c57d0b

Observation 8ffe0a2d-f610-42fa-841c-d5d9ab45293f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini: A Family of Highly Capable Multimodal Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:14.937333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:14.937333Z digest=sha256:15cdabbb7e555422e22f81fa293049429c775adbf62fc77b0fc22564a84fd07d

Observation 34992c8b-5c2f-461a-9ebd-97bd33f16688 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.034425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.034425Z digest=sha256:6bf9b50732edb38c056d16e93ce23025318476807cb02399dd7ecc8ba8820d58

Observation dfa936c2-1361-44c1-aca6-700a74758dfc · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini Robotics: Bringing AI into the Physical World

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.105152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.105152Z digest=sha256:73a99d6128aac0177f19957ac0c05067e47253947d2e3e7ee62cec7eb551eec0

Observation 61693193-d58e-4177-8a7d-ee10df47b7af · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaMA: Open and Efficient Foundation Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.176031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.176031Z digest=sha256:f5209b7f02769c922db1d2b57f165c5733221e593ded273b2ba451a154ac72b1

Observation 06719993-6dd7-4eac-95f8-55b4e8a21ec3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.268308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.268308Z digest=sha256:da55bc36d731c8ae535012a3b5cdb389c2ef1c4a0d51fc1bd20ffd04fc98a203

Observation a41ee9bc-cd25-455a-8f98-20c48cd04f0e · outbound

This paper cites The all-seeing project: Towards panoptic visual recognition and understanding of the open world.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces The all-seeing project: Towards panoptic visual recognition and understanding of the open world

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.827901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:15.352827Z digest=sha256:c3687c7302a5307ef300947e3657bc207db7f3bfaf3d3b56cf8887f8ba7a8967

Observation e1d03a9a-3438-4690-a6a9-9da73d0da9e6 · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.449497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.449497Z digest=sha256:ddeb139f40adb09c988619f76f536a05482953650e17ce0a417bc6f6307a12bb

Observation e412d259-4bc0-4a4d-8e91-d339d6ca3b4e · outbound

This paper cites Grok-1.5 vision preview.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Grok-1.5 vision preview

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:19.721585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:15.510411Z digest=sha256:f4f636e44d44838d92c5015ea80fc34c4bb3e85cb105b90d0f5eb9cfd5c8b9bf

Observation eaff1423-9048-4f1f-8ee2-dd20b974265f · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.576937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.576937Z digest=sha256:5bd0caa3658a8ef2f25b2b6bb8b00db0dc01430cd7ee46de874d6335ae17f8e3

Observation 239354fb-65f5-4647-9534-f23f80f943f4 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.641523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.641523Z digest=sha256:83e0900c48d13db473742d2b634461664998b44c5c837c38c95c36f1432bb6d5

Observation 8944cc7b-10b1-4c98-b1bf-5159dde0e1f4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:15.718768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:15.718768Z digest=sha256:9486941636396132f155ef95efcd63bcfed161b436d30f7fa496118c625f5a99

Pith citing papers

Observation 92387108-7980-4289-8454-b5ce4b3b9582 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.605186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.605186Z digest=sha256:f0d774b54a07fe91a33b036cb9180ae363c92828e06d6c5ae0081bfc927fa995

Observation 6253ce52-1ae4-49db-9f13-5e9949c949b3 · inbound

Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs cites this paper.

Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:29.157371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:37:38.901354Z digest=sha256:ae6357270c313a5051e492a059b5872b393f3f7e0634509208ed76ef3798a9f1

Observation cc80ffac-da6b-477b-bd31-f75cdccac3fc · inbound

The high-speed X-ray camera on AXIS: design and performance updates cites this paper.

The high-speed X-ray camera on AXIS: design and performance updates Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T18:48:03.418621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:48:03.418621Z digest=sha256:e23f8c4cd7e48b34d64345a8105894d4eb2072a7f2fead19e8d8adc6f0006b49

Observation 8684ed19-c732-44f0-b64a-6561ec1fe91b · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.932813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.932813Z digest=sha256:f23ebdbb3cb98a2671f486e9a9e0fb405e30d8c4e986f77c8070d2bdae5d9762

Observation 80a4f563-9ab0-4a4d-96c2-d8e8d525ba01 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:39.833644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:1f49d6acf335a39c6e3d9720c0f99d8301acce05e65aa43fb1845dda8588c8eb

Observation 35c55ac7-b8a4-4f5b-ab32-c4361e29ce64 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.714505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:8b380e28c65dec16b3a4fa036e04f4fcbf2a3b31388b0ef3c50207eb59c26e3c

Observation 189a4e30-35d9-47ed-9375-7da26e19f9ba · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.373521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:339b0d87637b2e9810e551692d17647648ed0dfdf0c5008b9f126243a42554cb

Observation a7ea5a71-deb2-4e97-9c2b-897a4aad0f53 · inbound

3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding cites this paper.

3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:56:00.254914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:51:52.063238Z digest=sha256:b637e245efd7d8b08de2f33371411d50b49f61fd3407a3e0226ea9f04da66885

Observation 0749cb87-7f03-44c0-bac1-e3af4955cb14 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:58:49.605881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T17:57:00.909897Z digest=sha256:b9ee41f7c2ffb7fc44632894f5e075660a6467aa27407b292ec04fa71ea571ef

Observation 37aa5277-49f8-4b04-8a43-0d65279cff87 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.582050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T19:02:05.125937Z digest=sha256:e775043e204389c13e8043b5db4c9bafe29bf70a39dbd530145661dd816e3555

Observation 319808e9-17f6-4ce6-90d6-8db16d291176 · inbound

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation cites this paper.

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:24:40.387049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:24:30.669956Z digest=sha256:082ca4abaf83c778c0f219188c89e996303adb4dad7169aab26596d113e3081f

Observation 63e8f605-6c9f-4107-bda1-eca7581de84b · inbound

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation cites this paper.

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.282353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:07:49.005419Z digest=sha256:2b3fba5c336af742e30136c6abec703dd84f812e3f1dafd55e18b10eb0ccb725

Observation f8731dac-224d-47d0-9833-45ab6014a0b7 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:58.950221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:c760ea7e4e76571479e4115e9e751610ad0314d9445b2af8f48c9e97f7cf84fd

Observation b4b13735-5d8f-4332-935e-4cce4db729cf · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.937015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:305e67ceb91b6dad1500d47183dd7acd381efba585013156a79e2335035e7c04

Observation cdae8602-7c06-4f94-b73c-6ab9c10f8ddd · inbound

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs cites this paper.

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.075845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:02:07.122615Z digest=sha256:20351a798381f384794e5212713a1e4f20141ebf633c9992a605172fb76ca94a

Observation 8b86d7ef-1904-4510-9c25-e011c7353b06 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.154508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:89d5152940e4712fa8161a18cec2fb71f6c3d4d6e7c289adf35c88f5c6a4cf03

Observation ad9c88a9-defe-4f09-9434-36ed67cab158 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:700a43e56c42ebb7a72c2a24fb928242f13c217fafe4e66173834f0549a1b13e

Observation be10cee9-94d9-40db-a399-91baf0cabd54 · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.760148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:26:32.209719Z digest=sha256:326c1b507fe1b2e75e9b70d06956fa91ebf45dcefe715162d3ab6bbdb0090464

Observation ca74376b-ee0f-4340-892e-ef46ed4806a6 · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T11:43:49.639057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:43:49.639057Z digest=sha256:d1160dab11d5cb1310941a69baaad8bce7512af246127f84034df97c1504fd6d

Observation e7ce9633-15b2-48d6-adad-d5b498358984 · inbound

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale cites this paper.

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.351453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:28:14.694168Z digest=sha256:490df4e803d8f578ce1524a2cdcc9b5080a4a717aeb0f01e99e8133813e4092a

Observation 880ac1a5-6849-4c80-8dfe-753e815bd6bb · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.366106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T11:17:48.279808Z digest=sha256:4cebfff7782baf153476c0e8fafe3b96ee7f9fc6ed5378f6c6ca909b89b3240d

Observation 9d610a94-38fe-468c-b91c-69837c2cc378 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T11:20:26.553333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:20:26.553333Z digest=sha256:223233012f08824206f97b172ee0ecd9822db8faadbb381cd50ee70bf741314d

Observation 5ca3a554-d397-484d-9748-597a412c8479 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:ab78a33591a755e9137eab37d54c23d137607bbb2d5b162fc0cf1f7b967737c8

Observation e0899ac6-ff62-4be6-bdbd-271e62038c2b · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:39.709300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:39.709300Z digest=sha256:32724d01980d003e0b6fa23d42193993f43e5ae982e7e5193400b6aeb82ef08c