Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:15.718768Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 24 inbound Pith citation observations for arXiv:2506.00123.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:15.718768Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:47:29.605186Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T15:28:34.757769Z
100 of 112 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1ee2d789-c94a-4f9b-a676-0cb0ed3ee942 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0674a851-2993-401b-9646-6fb3e9800e24 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8518c5f4-2e1c-4473-9f98-374b29f9bc36 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb545fed-1e29-4216-b41d-865137d9a172 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Claude 3.5 sonnet.https://www.anthropic.com/news/claude-3-5-sonnet, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e05a99f-c97b-4d81-8e45-9c47448f95e1 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc79de96-659e-4e20-85b5-93b6aabad814 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scanqa: 3d question answering for spatial scene understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e7458e-0d66-4212-a353-de4c5255c4d4 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Qwen Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d9a7022-2f3a-407a-8a0b-fd1d9156406f · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Qwen2.5-VL Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c32fc78d-0066-4baf-b491-4b4462ee1345 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38f2732e-a4a2-4351-b414-6069c7abdd9b · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RT-1: Robotics Transformer for Real-World Control at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0719468-b015-4571-b5ab-ce5ab0cee77c · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 696b5cb8-bf54-4657-aabd-a0663a8265c3 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language models are few-shot learners
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b5be8d8-f62e-46fc-830e-a9989623dece · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 105cfb30-dcb7-4e70-8984-d9e75ec5ad05 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8675d496-772e-46a3-bca9-59a9e49115ab · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Eagle 2.5: Boosting long-context post-training for frontier vision-language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e55fa0c-5216-4a00-976a-ae8b94f5868f · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4876270-6c15-403f-a4a6-07c8cb24c66d · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Sharegpt4v: Improving large multi-modal models with better captions
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1407922-9ad2-4d3f-a646-e4f4e64c2731 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3984c22a-1ba4-4e61-8e53-46502b5649fb · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf511ce-2e13-437b-8378-3041049502fe · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89df5c1a-f92d-4388-af89-4fa46b24b6bb · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38816cc3-5d58-4c66-9782-0b54dcd0d5d0 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces NaVILA: Legged Robot Vision-Language-Action Model for Navigation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75ef2d6e-407d-4877-85d3-9cfbbaa58328 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Spatially-Aware Transformer for Embodied Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80793765-2150-4de9-a05f-3f5f0f266ed7 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Local all-pair correspondence for point tracking
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d538287-bf9c-432d-833a-3fc1eeb760c6 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Simple and effective multi-paragraph reading comprehension
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d85376bb-fe9a-4889-b1b1-c7529397f013 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9cd709d-1052-43be-89e4-9d466b983348 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language modeling with gated convolutional networks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d053136-7bcb-43c0-81e8-7dbfb24ed116 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Quar-vla: Vision-language-action model for quadruped robots
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e06e275e-ba11-47ba-aa0b-f6fc31c2dc26 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces An image is worth 16x16 words: Transformers for image recognition at scale
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db3c8b1-01e1-4c50-9129-6fd6c48ccb2a · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces PaLM-E: An Embodied Multimodal Language Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a0ae299-5844-406f-b090-fe6419e2ce68 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Centernet: Keypoint triplets for object detection
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4629046-9ed0-43c9-80a3-4f6dfb2ca8a7 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8224738-d488-4f4f-8a18-9997eb7fa620 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Eva: Exploring the limits of masked visual representation learning at scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c64dda6-17a7-4488-b991-c933513eb539 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21a4e9b-e31b-4fac-89a6-3bdabb2ecf9c · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Rlafford: End-to-end affordance learning for robotic manipulation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49009cd-d63c-43e9-8140-b302563f7a46 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 442877b5-7d93-4dc8-9674-df4b5cbf5740 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee8501e-5fbc-4928-8892-5c0aed834f35 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Girshick
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d24edc9a-76f8-4ada-9938-76c194bc2301 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems , 36: 20482–20494, 2023
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ed7cf1-98a9-455c-92d6-50f5faa34813 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c4ad44-4a1e-4e28-a4ca-ca0c0ec860da · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chat-3d v2: Bridging 3d scene and large language models with object identifiers.CoRR, 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc41a651-4953-4a73-932e-78c486576128 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces An Embodied Generalist Agent in 3D World
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08a66b9-b8f2-4040-b0e6-1a299e52ec87 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Inner Monologue: Embodied Reasoning through Planning with Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834ff20c-66be-4c9f-b9db-6b92c79dadcc · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Bc-z: Zero-shot task generalization with robotic imitation learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e4ec0b-04f0-4385-a4a5-3a983426f210 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7dafc7f-ff80-4db9-b147-1bce094aaacd · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Scaling up visual and vision-language representation learning with noisy text supervision
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44b23ff9-3d8f-47d5-a5bf-9b174e55b26c · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces A diagram is worth a dozen images
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf0d064b-94e1-4314-b8cc-85003f6d3a3e · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OpenVLA: An Open-Source Vision-Language-Action Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f767ef-56a2-4155-a22e-5bbff6497ade · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LISA: Reasoning Segmentation via Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe47318a-98c5-404b-89ef-ab3a454e34fc · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Cornernet: Detecting objects as paired keypoints
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89609035-90f0-40cf-a3e8-4f2913f97ebc · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection.The International journal of robotics research, 37(4-5):421–436, 2018
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87c0ae9-e60d-4481-af81-960fba18264f · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb31d28c-6296-4572-9949-15950c5d4bab · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaVA-OneVision: Easy Visual Task Transfer
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d328658-9d72-4246-910b-1096cea908fd · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning agile skills via adversarial imitation of rough partial demonstrations
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96670f96-a2ab-47ea-be94-64bf0c89c8af · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3e139c2-7507-4bfa-abf1-f3de348530d5 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3524013-2008-4470-84d9-193b96830b34 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Exploring Plain Vision Transformer Backbones for Object Detection
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c551f63-1ff8-4ab4-af58-6f087d92ba59 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Robotic Visual Instruction
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f210c89-75f5-4c59-8221-675fa03e43a8 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Code as policies: Language model programs for embodied control
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f13fe24-c711-4692-bfd3-03c1f25cdbde · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Vila: On pre-training for visual language models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a3dc0b-fc1d-446d-ac2b-f7452a92d593 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06ca9977-97ca-4c6e-a229-b1cf9a2edc0a · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3daa9635-42f0-475b-8fd9-4acffa328232 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Visual instruction tuning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36f2259d-718a-4b25-8a9e-45b2f0ae47fc · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MMBench: Is Your Multi-modal Model an All-around Player?
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689a1580-0f37-41e1-a0c9-08676e08dfa2 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a306407-9e2e-46b0-b61d-48280076fd1e · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e31acdc-ec99-411c-8518-ebe57bced002 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef13d95-c7ca-4dad-b000-c89cc7399d84 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e872cd2-3e60-48b5-b16e-e9c2b767fd75 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SQA3D: Situated Question Answering in 3D Scenes
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3213c8d1-cdf6-4d3c-ad23-d6aa78d04de1 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems , 36:655–677, 2023
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0efa2d26-a987-441b-8041-4ffb0a5eeab3 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Walk these ways: Tuning robot control for generalization with multiplicity of behavior
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 060f98ff-3b53-4e2c-a68b-e3d12ac5f5be · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8129ca85-7683-47cb-b8ce-15aea4b9dbb6 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3251f2c2-46c5-4ac7-92ad-062f19267b9b · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Infograph- icvqa
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 453ef401-645e-4c86-b12e-a34b569110d1 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93073953-f24b-452a-846b-123bb4f6078a · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Embodiedgpt: Vision-language pre-training via embodied chain of thought.Advances in Neural Information Processing Systems, 36:25081–25094, 2023
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20ec6a30-082c-4364-9530-38aabc8e2478 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces R3M: A Universal Visual Representation for Robot Manipulation
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d906ba-aa13-49cb-bb2d-79eef539f927 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gpt-4o system card
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fa5edb8-db68-4971-9e85-c132ddab602c · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 859eb22a-392a-4190-b6e6-45f594123c8b · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0efc46c-7e1b-425c-8ae2-b32a44a583ca · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Improving language understanding by generative pre-training
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c89e3d8-d4fe-4810-833e-1931d3ffc11c · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Language models are unsupervised multitask learners
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee3335e7-0fd7-4dfd-b5a8-50e2e5a0eabb · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Learning transferable visual models from natural language supervision
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf7157ea-f21c-4296-8883-60c90dca772f · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Real-world robot learning with masked visual pre-training
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9add7871-7f36-4f34-a621-f36582eb566d · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Cliport: What and where pathways for robotic manipulation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a71a9fb-6724-476b-abb4-411f7b0e9b7c · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Towards VQA models that can read
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6059b5a-5b2a-4cda-8596-327f9ed638db · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Open-World Object Manipulation using Pre-trained Vision-Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74674bf9-e778-4018-b380-b4b587df19cd · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SMART: Self-supervised Multi-task pretrAining with contRol Transformers
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffe0a2d-f610-42fa-841c-d5d9ab45293f · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini: A Family of Highly Capable Multimodal Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34992c8b-5c2f-461a-9ebd-97bd33f16688 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa936c2-1361-44c1-aca6-700a74758dfc · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Gemini Robotics: Bringing AI into the Physical World
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61693193-d58e-4177-8a7d-ee10df47b7af · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaMA: Open and Efficient Foundation Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06719993-6dd7-4eac-95f8-55b4e8a21ec3 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a41ee9bc-cd25-455a-8f98-20c48cd04f0e · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces The all-seeing project: Towards panoptic visual recognition and understanding of the open world
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1d03a9a-3438-4690-a6a9-9da73d0da9e6 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e412d259-4bc0-4a4d-8e91-d339d6ca3b4e · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Grok-1.5 vision preview
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eaff1423-9048-4f1f-8ee2-dd20b974265f · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 239354fb-65f5-4647-9534-f23f80f943f4 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8944cc7b-10b1-4c98-b1bf-5159dde0e1f4 · outbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92387108-7980-4289-8454-b5ce4b3b9582 · inbound
RoboBrain 2.0 Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6253ce52-1ae4-49db-9f13-5e9949c949b3 · inbound
Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc80ffac-da6b-477b-bd31-f75cdccac3fc · inbound
The high-speed X-ray camera on AXIS: design and performance updates Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8684ed19-c732-44f0-b64a-6561ec1fe91b · inbound
Contrastive Representation Regularization for Vision-Language-Action Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a4f563-9ab0-4a4d-96c2-d8e8d525ba01 · inbound
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 35c55ac7-b8a4-4f5b-ab32-c4361e29ce64 · inbound
MiMo-Embodied: X-Embodied Foundation Model Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 189a4e30-35d9-47ed-9375-7da26e19f9ba · inbound
Token Warping Helps MLLMs Look from Nearby Viewpoints Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7ea5a71-deb2-4e97-9c2b-897a4aad0f53 · inbound
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0749cb87-7f03-44c0-bac1-e3af4955cb14 · inbound
GeoWorld-VLM: Geometry from World Models for Vision-Language Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37aa5277-49f8-4b04-8a43-0d65279cff87 · inbound
GeoWorld-VLM: Geometry from World Models for Vision-Language Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 319808e9-17f6-4ce6-90d6-8db16d291176 · inbound
SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63e8f605-6c9f-4107-bda1-eca7581de84b · inbound
SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8731dac-224d-47d0-9833-45ab6014a0b7 · inbound
Extending Embodied Question Answering from Perception to Decision Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4b13735-5d8f-4332-935e-4cce4db729cf · inbound
GEM: Generative Supervision Helps Embodied Intelligence Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdae8602-7c06-4f94-b73c-6ab9c10f8ddd · inbound
OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b86d7ef-1904-4510-9c25-e011c7353b06 · inbound
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad9c88a9-defe-4f09-9434-36ed67cab158 · inbound
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be10cee9-94d9-40db-a399-91baf0cabd54 · inbound
RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca74376b-ee0f-4340-892e-ef46ed4806a6 · inbound
RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7ce9633-15b2-48d6-adad-d5b498358984 · inbound
SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 880ac1a5-6849-4c80-8dfe-753e815bd6bb · inbound
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d610a94-38fe-468c-b91c-69837c2cc378 · inbound
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca3a554-d397-484d-9748-597a412c8479 · inbound
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0899ac6-ff62-4be6-bdbd-271e62038c2b · inbound
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.