Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:14.304184Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 69 inbound Pith citation observations for arXiv:2506.03569.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:14.304184Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:22:55.993691Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:29:57.282206Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9aa704c6-9745-4a83-9db6-73fd870606d3 · outbound
MiMo-VL Technical Report Alayrac, J
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2826d57-f206-4ba1-be60-983d963a0685 · outbound
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50a1dfbd-b5e8-4c7c-a3a4-32c2cb1cd12c · outbound
MiMo-VL Technical Report $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 670331bf-23e3-4d08-8d1e-a9d3fb28f477 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ef05f39-56ca-4db7-97e5-7c698381868f · outbound
MiMo-VL Technical Report WebSRC: A Dataset for Web-Based Structural Reading Comprehension
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106c8006-b943-42f6-89be-79b3fc2895ba · outbound
MiMo-VL Technical Report AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6488aff5-0583-4f75-86e2-4a7535c09af9 · outbound
MiMo-VL Technical Report SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff283ef-5fb7-446c-934f-e085a2eb0250 · outbound
MiMo-VL Technical Report VisionArena: 230K Real World User-VLM Conversations with Preference Labels
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1f6d652-58c4-4a40-8837-16795209b209 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c1e6125-b66b-4986-9ca1-38897496c9e4 · outbound
MiMo-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8798a9b4-020a-40fc-8455-e5210490734d · outbound
MiMo-VL Technical Report SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 298525a3-dc3f-44c8-bb97-016218fff5ec · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc824662-9735-4f92-aa61-5b0ab13b7362 · outbound
MiMo-VL Technical Report Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c985dc53-33c1-440a-945b-883a04ae9c81 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4c3bb18-9b57-45d0-bfea-258c3bcfb8ac · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747b7261-a0d9-4748-8ab5-e02023921a45 · outbound
MiMo-VL Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 987da426-a6e2-4e58-89e8-0e770f4d570b · outbound
MiMo-VL Technical Report Measuring Mathematical Problem Solving With the MATH Dataset
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d211c42b-33cd-43bf-a248-ebabe66b7dc9 · outbound
MiMo-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483754e1-be5e-4ab1-bffb-67493c08e6f3 · outbound
MiMo-VL Technical Report MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c2bb9ce-966a-4916-bb14-8d66bfb6ff6c · outbound
MiMo-VL Technical Report Karamcheti, S
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b54d876d-a504-455c-a503-e2fdd9f8b944 · outbound
MiMo-VL Technical Report Kazemzadeh, V
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38e4f808-d3be-46c9-936a-82138f84335e · outbound
MiMo-VL Technical Report Kembhavi, M
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da410722-19b9-450b-b315-c59a0296f470 · outbound
MiMo-VL Technical Report Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24a4638-6649-4dad-9272-65e048d80726 · outbound
MiMo-VL Technical Report SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a081a7-53ea-40d3-a02d-6246af5351d4 · outbound
MiMo-VL Technical Report ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 350c9168-2d72-4d6f-92d7-aba2d8543642 · outbound
MiMo-VL Technical Report VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e454d1-eab0-42c5-9876-16cbc6fc0224 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db0c7814-3e4a-4b6f-86b6-8e5a0d247c6d · outbound
MiMo-VL Technical Report VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b776c9-bfe4-4024-868a-d8d1b4f460a9 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50380670-97b8-45ca-b403-70f4ae033c09 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c7eed2d-cf0e-41f6-8551-7053f830d656 · outbound
MiMo-VL Technical Report MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef01348-6663-492a-9c0e-ae9a7d88f4e0 · outbound
MiMo-VL Technical Report American invitational mathematics examination - aime
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ce0a271-881f-47b4-8ed1-e218205e07d2 · outbound
MiMo-VL Technical Report American invitational mathematics examination - aime
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05505eba-d9a1-42e4-8397-5fd8b2c0c4f1 · outbound
MiMo-VL Technical Report Mangalam, R
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 672eaaeb-cfdd-4f7b-8ac2-88fcf873d575 · outbound
MiMo-VL Technical Report ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7409a732-62cd-4b2c-ad5f-db677b1d8397 · outbound
MiMo-VL Technical Report Mathew, D
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d41bc7b-b4d2-4d45-9f17-79c5146d37ac · outbound
MiMo-VL Technical Report Mathew, V
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0b5c300-4c4b-4f15-b85b-fa571334f0b2 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9e6d50f-7a39-4fd0-a161-ca5687bd38fa · outbound
MiMo-VL Technical Report Computer-using agent: Introducing a universal interface for ai to interact with the digital world
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38547bc7-66bb-49a5-973d-4edba5f007c4 · outbound
MiMo-VL Technical Report Ouyang, J
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 393c8537-f874-4d32-8226-f091b85810bc · outbound
MiMo-VL Technical Report Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dfcd675-3d31-4dbf-b1b1-035ce52f7aa0 · outbound
MiMo-VL Technical Report Paiss, A
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f93ff15b-64ec-4967-b269-e728eda1e8f6 · outbound
MiMo-VL Technical Report We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a32917-7a98-4a22-8e4f-fd16a6b7521b · outbound
MiMo-VL Technical Report UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7c6114-3b17-4b08-a388-ef2a0bdc0a10 · outbound
MiMo-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a87f416-d202-47dd-99c8-72f7bfeed1cb · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 194075af-6572-413f-9fb6-d5da771e53de · outbound
MiMo-VL Technical Report Rezatofighi, N
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7c35073-bd3c-4415-9f5e-2b0e483d572b · outbound
MiMo-VL Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1886f12-039c-49d5-9267-da3c22794b7c · outbound
MiMo-VL Technical Report HybridFlow: A Flexible and Efficient RLHF Framework
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d03c616-b330-429d-bf9b-6b6a85f18dd8 · outbound
MiMo-VL Technical Report Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3fcf5b8-5082-45d0-836b-ec5d85db990a · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a307c3cd-758c-45b7-a4c2-7b57ef8323ba · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5df548b-aabe-43e9-9ea2-31398e8463c5 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12058bba-b8d3-4e0e-bbc5-eab30b9b953e · outbound
MiMo-VL Technical Report Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e708e028-114c-4a3e-b3dd-ed2805a5e282 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7599ee8d-24d9-4414-ba22-ad3aa2d74bf5 · outbound
MiMo-VL Technical Report V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9641ad7-43ab-4fd8-919f-7e04dbdcd212 · outbound
MiMo-VL Technical Report OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a3932f-3008-446c-b213-fbec0feedcc5 · outbound
MiMo-VL Technical Report LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64865b0b-3a69-4a7f-9f3a-9ebbb71b8cf6 · outbound
MiMo-VL Technical Report MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0557dbb1-cc66-48fa-902e-f0342f4081c9 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f58e01f-8ddc-42bf-84f5-462322cdc782 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41847435-2c5a-495e-8d21-d61fea1e607a · outbound
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c68946e-d713-4c77-a7f8-b444e217f206 · outbound
MiMo-VL Technical Report Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ae787b-39e1-450b-a93d-7d962954c39d · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d02eabf1-fe77-4d49-82fa-7cc796e5fc9d · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03125271-f4c1-45c0-89b3-0ed29a95ad57 · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3befae66-651f-4867-b1e4-fa7177b06cfd · outbound
MiMo-VL Technical Report Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation feaed296-f60b-4e70-9790-c248f2ad0358 · outbound
MiMo-VL Technical Report MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399f54ee-1809-460e-a0d9-df04827b9876 · outbound
MiMo-VL Technical Report LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ef15d8-c794-4f89-ab02-c0dde79acb97 · outbound
MiMo-VL Technical Report MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d2026df-3505-4a46-96e4-998670e66d1d · outbound
MiMo-VL Technical Report MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67daf7b8-965b-439a-9e86-0d81d6023e52 · outbound
MiMo-VL Technical Report Zheng, W
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ecb9df86-496c-40c5-ad32-5a02543195af · outbound
MiMo-VL Technical Report Instruction-Following Evaluation for Large Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7578a60e-6a0d-4f9f-9c08-77af19bdbbc3 · outbound
MiMo-VL Technical Report Zitkovich, T
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b61006d-4166-4d49-ad64-4c018a689bad · outbound
MiMo-VL Technical Report DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28b2eff-3c30-4dc7-84c2-4f4f91a8c0bb · inbound
Reinforcement Learning from Human Feedback MiMo-VL Technical Report
Reference 174
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61343ccd-0658-4382-8a41-57d7baedba20 · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MiMo-VL Technical Report
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c8541a-7675-4cd5-9fc8-2cb2f4bb202a · inbound
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 643527ce-67f8-431a-85b6-f4d5b2f11e23 · inbound
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e0d4a7-b639-4306-8547-51f65d2169d2 · inbound
How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study MiMo-VL Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf116d3b-3523-42a2-b6d4-71f203e48e34 · inbound
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents MiMo-VL Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1b8f95-21da-41f6-abcb-9ef99a3cb01b · inbound
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback MiMo-VL Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d8ff12c-2c07-4a63-99f6-b547ba3544cc · inbound
EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO MiMo-VL Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71aee071-de03-4200-b863-2ceaaa125bea · inbound
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models MiMo-VL Technical Report
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd128e0-de73-428a-b81b-8ee18a9c45cc · inbound
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning MiMo-VL Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b3a079a-698a-45a6-aa9f-76a8f375d4b1 · inbound
Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models MiMo-VL Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318485fe-dbb2-4117-bac5-cd3e05544081 · inbound
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning MiMo-VL Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1283b163-a25e-4989-99da-ea011bde4706 · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MiMo-VL Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f477d4-a0ee-4e46-b838-1ef3770b3e50 · inbound
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66398f4f-8c4e-4ba3-b31d-33a74d5d8454 · inbound
Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing MiMo-VL Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c8d68e1-e4a3-49ab-ae44-cdf495f6c7a2 · inbound
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiMo-VL Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f71a09d5-6d54-4b7c-8b97-ee11e5773ca8 · inbound
VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation MiMo-VL Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297e34fa-3fdd-4af5-ade9-56901d741640 · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models MiMo-VL Technical Report
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b12fa6d3-9b00-4db2-95ce-68799fe4fb10 · inbound
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning MiMo-VL Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6250233-8598-4bfb-9014-55b6ec3f5792 · inbound
MiMo-Embodied: X-Embodied Foundation Model Technical Report MiMo-VL Technical Report
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c82b5776-93cb-4903-8d1a-afe3ddaf8a8c · inbound
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation MiMo-VL Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73ec2059-50bf-408b-a4d9-affbcfca6975 · inbound
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation MiMo-VL Technical Report
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd98a5d-fed8-4758-8c26-3f7e5d93553c · inbound
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? MiMo-VL Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8628686b-b34e-440d-8232-120cf27510e9 · inbound
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? MiMo-VL Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f537819-a78b-4e4e-8494-e6c407fc8e08 · inbound
Learning Self-Correction in Vision-Language Models via Rollout Augmentation MiMo-VL Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a820de4-a910-4a8d-a76c-6440165bd6fd · inbound
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MiMo-VL Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54be8578-9b88-423a-8e54-6368f2bb3019 · inbound
Visual Preference Optimization with Rubric Rewards MiMo-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ad5634b-2940-4275-8fe7-d9223b08657b · inbound
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations MiMo-VL Technical Report
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93ddcea5-acf8-4ffb-8237-018a3f76e2a3 · inbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MiMo-VL Technical Report
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccbc11e3-b927-4d3a-8df1-943a753f759e · inbound
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model MiMo-VL Technical Report
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0dd9a59-629e-4a98-aa0a-89287382ae77 · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MiMo-VL Technical Report
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77cabe27-51d5-40c4-aa38-b3a2361805c2 · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MiMo-VL Technical Report
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4354bf9-2a8f-4871-91f3-eb6334a7b9e1 · inbound
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning MiMo-VL Technical Report
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f5b740a-da5a-4aff-b708-a435da384cba · inbound
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos MiMo-VL Technical Report
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c894e30b-280e-44aa-b709-a1bfb5edb5a9 · inbound
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models MiMo-VL Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 893cbb9f-1030-4cab-b4cc-941550d89e86 · inbound
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ac97830-2507-4a0b-98ad-ce6f04e24b6e · inbound
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0569c09f-ce8d-4b67-80c0-7a74e654719c · inbound
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 031d66aa-8f0a-4a00-b416-864604493087 · inbound
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0334a5b5-72d2-44bf-8c11-5bb24cc56331 · inbound
Video-Zero: Self-Evolution Video Understanding MiMo-VL Technical Report
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e181396-0d43-4b55-809d-c86d586c3484 · inbound
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation MiMo-VL Technical Report
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9a0c321-523d-4464-b0a9-c3eeeeffb87a · inbound
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation MiMo-VL Technical Report
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f91bb6c9-f935-4c34-ab58-92281d03fbf3 · inbound
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues MiMo-VL Technical Report
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83e8f609-e1d0-4b50-9f14-5bc4fcd3faf0 · inbound
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding MiMo-VL Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7fa44757-a829-4dcb-9436-49bcdd706b9d · inbound
FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis MiMo-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5a691c8-6b18-4d0a-91c3-a657e0f751e0 · inbound
Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker MiMo-VL Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6ec13b3-b844-40b6-8ed8-9bcf9aa26393 · inbound
Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning MiMo-VL Technical Report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3b539c2-eae4-444f-ae51-e861be05b60f · inbound
TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL MiMo-VL Technical Report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af881701-d82d-4273-839e-3ef5afd8e92e · inbound
Benchmarking Visual State Tracking in Multimodal Video Understanding MiMo-VL Technical Report
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ba9fe9f-ca95-473a-a74f-74f5ed8b7c7e · inbound
Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues MiMo-VL Technical Report
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 229438a3-c704-4337-85dd-68bfdd291a7a · inbound
When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding MiMo-VL Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f312811b-cd21-4be4-9a69-dd19e9c8f32e · inbound
Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding MiMo-VL Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8f54eff-88c8-4653-95b8-f558789e19cf · inbound
AIR: Adaptive Interleaved Reasoning with Code in MLLMs MiMo-VL Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a3f5543-80b6-4747-bf82-9c2e3d3ef69f · inbound
Latent Visual States for Efficient Multimodal Reasoning MiMo-VL Technical Report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 991bc27b-06fc-4768-8e66-36a67bbedc2e · inbound
Aloe-Vision: Robust Vision-Language Models for Healthcare MiMo-VL Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36501310-a408-4f0d-89a1-e547603e6e68 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models MiMo-VL Technical Report
Reference 277
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 485ee3cd-1f75-48bb-ade0-16c332cb0008 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models MiMo-VL Technical Report
Reference 277
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3295a0c-7ae7-45dd-8c0d-f85951fb995d · inbound
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension MiMo-VL Technical Report
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45c10caa-2eb0-4a62-b485-f74be4e92864 · inbound
Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools MiMo-VL Technical Report
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb177ae-653c-4280-b4d2-86794a0e404f · inbound
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence MiMo-VL Technical Report
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c9a65e3-14b0-4b3d-bf4e-54edb7612293 · inbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs MiMo-VL Technical Report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc403541-2b97-4b33-9b74-c58a6aa4d875 · inbound
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement MiMo-VL Technical Report
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a6600d-e2b8-4c86-bd0e-470aeb0b9fee · inbound
PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis MiMo-VL Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72382549-ea31-40eb-a9bf-70cb567513b8 · inbound
NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model MiMo-VL Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 958d1bf7-dbaf-43b6-8f45-47cd133a3c08 · inbound
HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research MiMo-VL Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaca9b6d-975b-4fba-a7c6-4d74fcfd061d · inbound
MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes MiMo-VL Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a12be9-d250-4a05-8de4-ccfca55f846b · inbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning MiMo-VL Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0ff46f-77e5-48d1-af0e-05b385705b8f · inbound
CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding MiMo-VL Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d883eb-ec4d-4405-ad37-cf7b0f03d325 · inbound
OPD-V: Visual On-Policy Self-Distillation with Modality Balance MiMo-VL Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.