Pith. sign in

Paper Citation Record · LEDGER

MiMo-VL Technical Report

As of 10 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 71 inbound Pith citation observations for arXiv:2506.03569.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03569 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:14.304184Z

measured 146 of 146 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 71 of 71 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:59:13.756188Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:57.282206Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9aa704c6-9745-4a83-9db6-73fd870606d3 · outbound

This paper cites Alayrac, J.

MiMo-VL Technical Report Alayrac, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.116017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.025617Z digest=sha256:05238d0d907890898010be3c8ef2955ec426e90ac30e0cfa161790a0fc7dd5dd

Observation a2826d57-f206-4ba1-be60-983d963a0685 · outbound

This paper cites Qwen2.5-VL Technical Report.

MiMo-VL Technical Report Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.034077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.034077Z digest=sha256:fb3baa06cb0ed5fb653486428297ea0b19d0932fc638972f2b779648abdf7016

Observation 50a1dfbd-b5e8-4c7c-a3a4-32c2cb1cd12c · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

MiMo-VL Technical Report $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.038069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.038069Z digest=sha256:6fdbee1c8b238f177b2d60bb7919629370f0a123c4148479e5566797d9ffe3a1

Observation 670331bf-23e3-4d08-8d1e-a9d3fb28f477 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.105797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.042192Z digest=sha256:4f5f0d4ba55cc1c7acc1360635c457a3cac2ed1391aba69cb7af5eafd009fcca

Observation 3ef05f39-56ca-4db7-97e5-7c698381868f · outbound

This paper cites WebSRC: A Dataset for Web-Based Structural Reading Comprehension.

MiMo-VL Technical Report WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.045891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.045891Z digest=sha256:4c0f9ad1cca629092b2034c550e4860f98e5225fa95485729acd24b2a71d26ee

Observation 106c8006-b943-42f6-89be-79b3fc2895ba · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

MiMo-VL Technical Report AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.050518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.050518Z digest=sha256:0554eb690708fa1399b0ac0cdb18b520982303746ee5c33b1cc82f5b4e853018

Observation 6488aff5-0583-4f75-86e2-4a7535c09af9 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

MiMo-VL Technical Report SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.055373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.055373Z digest=sha256:f64e2558ceaa496eb110a9891ea371799d1931794ce5e25e8c296a4a4791ddc6

Observation 6ff283ef-5fb7-446c-934f-e085a2eb0250 · outbound

This paper cites VisionArena: 230K Real World User-VLM Conversations with Preference Labels.

MiMo-VL Technical Report VisionArena: 230K Real World User-VLM Conversations with Preference Labels

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:04:14.747082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.059734Z digest=sha256:d504e6f7527883d6478195f2d63737c2b711b829226b6ac82330e3180449a848

Observation d1f6d652-58c4-4a40-8837-16795209b209 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.063690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.063690Z digest=sha256:89a6c65cdbfe083a2b4af9f5658cc7dccd8bd65b00eb95a556794a2e49062e8a

Observation 4c1e6125-b66b-4986-9ca1-38897496c9e4 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

MiMo-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.067321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.067321Z digest=sha256:c95fc3925c763de1e2b66a4e0197a98bba4ec20669cfa53f813c9656ab4df24a

Observation 8798a9b4-020a-40fc-8455-e5210490734d · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

MiMo-VL Technical Report SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.070645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.070645Z digest=sha256:679de678b977d8694597c358600c1c0dad21c8d99eaf6df630987ba310fcc9b5

Observation 298525a3-dc3f-44c8-bb97-016218fff5ec · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 13

Resolution
verified exact
doi, observed 2026-08-07T11:04:14.336562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.074224Z digest=sha256:63d740590cb7d388a56bb69add4960becb45fb0f02022c391506dd574a4b0b5d

Observation dc824662-9735-4f92-aa61-5b0ab13b7362 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MiMo-VL Technical Report Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.077563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.077563Z digest=sha256:06d83320dcf4cc229ca246e531339228c7a012cc3834663af60f4fa8fbce7278

Observation c985dc53-33c1-440a-945b-883a04ae9c81 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.090052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.081196Z digest=sha256:188c79d480bca42d0ef41ef8ba2b67931b2ea908fb80de1d975fb2d6c7f3e07f

Observation a4c3bb18-9b57-45d0-bfea-258c3bcfb8ac · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.085568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.085568Z digest=sha256:89f15615c3f590be938d76f4fe108f2f2abb7001c20693421667d3f81c396947

Observation 747b7261-a0d9-4748-8ab5-e02023921a45 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

MiMo-VL Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.088974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.088974Z digest=sha256:27bf92d1469cd652032e12ba93c709738f363d395f3fa523d94f9370f9c61708

Observation 987da426-a6e2-4e58-89e8-0e770f4d570b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MiMo-VL Technical Report Measuring Mathematical Problem Solving With the MATH Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.092481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.092481Z digest=sha256:95529f66ee3f202689f31a3daaa43e5a2115f5fa09aa022ef88ac24f0b0841c1

Observation d211c42b-33cd-43bf-a248-ebabe66b7dc9 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

MiMo-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.095997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.095997Z digest=sha256:93693a1cf7a0535936fcff6832e485b34d462ed281f14bb121b124b7644f5204

Observation 483754e1-be5e-4ab1-bffb-67493c08e6f3 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

MiMo-VL Technical Report MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.099478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.099478Z digest=sha256:67d4136bc175cf7847d6fb04d1bddcb962b8b8d76457db63aa7f4cde8ad2c741

Observation 4c2bb9ce-966a-4916-bb14-8d66bfb6ff6c · outbound

This paper cites Karamcheti, S.

MiMo-VL Technical Report Karamcheti, S

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.073290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.103006Z digest=sha256:9d01f49fc60e02c6024f2b4bc4570aa7bc3ca5fe39089ead68a9f7cd87fa4ef8

Observation b54d876d-a504-455c-a503-e2fdd9f8b944 · outbound

This paper cites Kazemzadeh, V.

MiMo-VL Technical Report Kazemzadeh, V

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.062551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.106351Z digest=sha256:1a7c5fa96321e92c145ccbf5f6dfbf83405af39bf5e7bf0976c8cb9a762bf184

Observation 38e4f808-d3be-46c9-936a-82138f84335e · outbound

This paper cites Kembhavi, M.

MiMo-VL Technical Report Kembhavi, M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.052054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.109684Z digest=sha256:b9bc8d0fffdd10a6476d2bb310f5b499eeefb181dec91794196f99c589a9db15

Observation da410722-19b9-450b-b315-c59a0296f470 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

MiMo-VL Technical Report Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.112913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.112913Z digest=sha256:8df9ec314f37e2d929c4be4cdaf87470449b800c9a3aef108d213725f5c2e3cc

Observation f24a4638-6649-4dad-9272-65e048d80726 · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

MiMo-VL Technical Report SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.116722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.116722Z digest=sha256:10a9224d0968d57263a528ab44fe0b4820dba246e9c127a6efc02120c898f5e4

Observation 05a081a7-53ea-40d3-a02d-6246af5351d4 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

MiMo-VL Technical Report ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.120139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.120139Z digest=sha256:7af3c04461290128752ef2dec88b78e7b6b56969d52884ca24f6b3f74592559f

Observation 350c9168-2d72-4d6f-92d7-aba2d8543642 · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

MiMo-VL Technical Report VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.123685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.123685Z digest=sha256:67a4a3ec776d3e9fb95799e5b0b6ec69abeb9012b1560f173084524f76626ca0

Observation 77e454d1-eab0-42c5-9876-16cbc6fc0224 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.040120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.127539Z digest=sha256:2ee7b639dcc6db6ff89fbc3cb7307f170f2a375c6dcbd39feecff50cc5465076

Observation db0c7814-3e4a-4b6f-86b6-8e5a0d247c6d · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

MiMo-VL Technical Report VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.131074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.131074Z digest=sha256:417f19777dad1d885534bf908d4c1b591f3d407fcaf506afb6b61b158afb30e5

Observation c8b776c9-bfe4-4024-868a-d8d1b4f460a9 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.030409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.135029Z digest=sha256:6f00325fffe5330b8417d1416d04b5058a3b6bb9b6370230459f1a779431b7b7

Observation 50380670-97b8-45ca-b403-70f4ae033c09 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.020268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.138906Z digest=sha256:43946a9c187aab1324dc9fa9e1d150b561b67b4d14848690fb40cfc55408ea5e

Observation 6c7eed2d-cf0e-41f6-8551-7053f830d656 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MiMo-VL Technical Report MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.142694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.142694Z digest=sha256:40ec5489cecbcc949ccafe6689ed1e9ec053569495f0b7981071208a77ed10d0

Observation 4ef01348-6663-492a-9c0e-ae9a7d88f4e0 · outbound

This paper cites American invitational mathematics examination - aime.

MiMo-VL Technical Report American invitational mathematics examination - aime

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.009244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.148173Z digest=sha256:2cb1852d62395d50abeb500953d5fc43ac9023e319b50dc78fc4413edf3fb77a

Observation 5ce0a271-881f-47b4-8ed1-e218205e07d2 · outbound

This paper cites American invitational mathematics examination - aime.

MiMo-VL Technical Report American invitational mathematics examination - aime

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.997894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.151937Z digest=sha256:ce71be59efb4ea51a1661c04b9b6c5985b05ef43aebce602f047e4bbe4e4a4aa

Observation 05505eba-d9a1-42e4-8397-5fd8b2c0c4f1 · outbound

This paper cites Mangalam, R.

MiMo-VL Technical Report Mangalam, R

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.988056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.156916Z digest=sha256:b37d6262400dc3e3ee7c8c7f192ca36f1cc18e369b6ecca876dbb4bd02cea243

Observation 672eaaeb-cfdd-4f7b-8ac2-88fcf873d575 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MiMo-VL Technical Report ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.160624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.160624Z digest=sha256:529397b492568598ad43cc17e20640df3894d2c4b14006c74348d9c258215fd9

Observation 7409a732-62cd-4b2c-ad5f-db677b1d8397 · outbound

This paper cites Mathew, D.

MiMo-VL Technical Report Mathew, D

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.978055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.164676Z digest=sha256:0e837545810ab43043c47f3959db490c123959e30fb8f43de4538813b117a53d

Observation 0d41bc7b-b4d2-4d45-9f17-79c5146d37ac · outbound

This paper cites Mathew, V.

MiMo-VL Technical Report Mathew, V

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.967115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.168315Z digest=sha256:fcb27f82f588e3f2a7e3041fb52940e5d1df303094320fd5c604096770e0088a

Observation d0b5c300-4c4b-4f15-b85b-fa571334f0b2 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.957073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.174943Z digest=sha256:11a02fb29db35a6a48cac7f4f8c5f390d060f73b1d864b6488de74fd4e11a63d

Observation d9e6d50f-7a39-4fd0-a161-ca5687bd38fa · outbound

This paper cites Computer-using agent: Introducing a universal interface for ai to interact with the digital world.

MiMo-VL Technical Report Computer-using agent: Introducing a universal interface for ai to interact with the digital world

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.179094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.179094Z digest=sha256:638e1385982f0df42f9e98556b523b4684692078677176e2249eac514710fbf2

Observation 38547bc7-66bb-49a5-973d-4edba5f007c4 · outbound

This paper cites Ouyang, J.

MiMo-VL Technical Report Ouyang, J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.941293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.182835Z digest=sha256:b407c020562c9f78ec9d6691c2e2ce1706dbb0c0704be01c5303d319a5a3a3ae

Observation 393c8537-f874-4d32-8226-f091b85810bc · outbound

This paper cites Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models.

MiMo-VL Technical Report Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.186518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.186518Z digest=sha256:835da3a27b23a45d2656c1579169d4ef67c7bb5dd9954ee70bde23c1e61907e2

Observation 1dfcd675-3d31-4dbf-b1b1-035ce52f7aa0 · outbound

This paper cites Paiss, A.

MiMo-VL Technical Report Paiss, A

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.930841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.190120Z digest=sha256:ddc52ff07b4c98ad60213a24d8cf83baac466ccbae0efab4e1382a6882aaff32

Observation f93ff15b-64ec-4967-b269-e728eda1e8f6 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

MiMo-VL Technical Report We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.193506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.193506Z digest=sha256:aae38665b3cb979191c490d9799fe86380f17865e2d7ed7d35a0234118e02642

Observation 20a32917-7a98-4a22-8e4f-fd16a6b7521b · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

MiMo-VL Technical Report UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.200543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.200543Z digest=sha256:6210be03eb046c3d8c75b669deeceda4b37bdc2fd1388168d7383dd49d7ccf60

Observation 4a7c6114-3b17-4b08-a388-ef2a0bdc0a10 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

MiMo-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.204341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.204341Z digest=sha256:9cad0987671f011a214c39b43d8be2c7862856987f9a305e4e6d2138a2c0dfad

Observation 1a87f416-d202-47dd-99c8-72f7bfeed1cb · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.208120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.208120Z digest=sha256:2a9ffc45d96342c83d2ae2f8b7e205828c22320e1b4f6c9187df26ced18898de

Observation 194075af-6572-413f-9fb6-d5da771e53de · outbound

This paper cites Rezatofighi, N.

MiMo-VL Technical Report Rezatofighi, N

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.912787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.211431Z digest=sha256:f60508ebd1a1ec9a6a59fd3e1dda86af1aaa00710994fb618469ed1742adad2f

Observation f7c35073-bd3c-4415-9f5e-2b0e483d572b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MiMo-VL Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.214574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.214574Z digest=sha256:d3c48dfffd00785ad4b85ac327182a01e38ab50306cb51cea11e3c0edd8eb976

Observation d1886f12-039c-49d5-9267-da3c22794b7c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

MiMo-VL Technical Report HybridFlow: A Flexible and Efficient RLHF Framework

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.217839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.217839Z digest=sha256:15eff3fe78e127bf1ceb4cd76a76775333c3fb433192adeb2a7c8cc6c098f216

Observation 8d03c616-b330-429d-bf9b-6b6a85f18dd8 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MiMo-VL Technical Report Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.221120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.221120Z digest=sha256:d733a830e576f99d4e0f52553054bdd9c71a070ce2fed0d874658bac30c87e5f

Observation a3fcf5b8-5082-45d0-836b-ec5d85db990a · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.902472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.225004Z digest=sha256:02b46bd94026bf05211792e0bb0ffa22bce044eb832e6507c39e639beb0ef616

Observation a307c3cd-758c-45b7-a4c2-7b57ef8323ba · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.891258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.228284Z digest=sha256:0187a1f0eba39cc3a2db934dbad96f3a803b08c6d27d414b71f0faa46acf8f16

Observation f5df548b-aabe-43e9-9ea2-31398e8463c5 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.879951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.231882Z digest=sha256:e35904a3dd1c9a3d1142e16c5d6f523ecc15b54f3f189db69dd5196904d73c38

Observation 12058bba-b8d3-4e0e-bbc5-eab30b9b953e · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

MiMo-VL Technical Report Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.234976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.234976Z digest=sha256:e791e64ffd7623908d4dca74d5d22c8cc661b61a19170b289c3536cdbd56ba07

Observation e708e028-114c-4a3e-b3dd-ed2805a5e282 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.869525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.238436Z digest=sha256:bca7ac3242b638b8d66706f7b9405ecc56f4a700c4616a93e5bac215628a9c50

Observation 7599ee8d-24d9-4414-ba22-ad3aa2d74bf5 · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

MiMo-VL Technical Report V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.241662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.241662Z digest=sha256:400097920c8ad18fbcefd9c0d8fbb9c73e82376e7db674e2ef6bfcfd4be9b642

Observation c9641ad7-43ab-4fd8-919f-7e04dbdcd212 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

MiMo-VL Technical Report OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.245410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.245410Z digest=sha256:c5ead7c697793795d465631f870b71b5c69478a6cdefd8c9d683c0e5ca11496c

Observation 80a3932f-3008-446c-b213-fbec0feedcc5 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

MiMo-VL Technical Report LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.248887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.248887Z digest=sha256:9dbf9ca2386961ebc19e2268199a8ca116655787b858ed702772af39c5814a22

Observation 64865b0b-3a69-4a7f-9f3a-9ebbb71b8cf6 · outbound

This paper cites MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining.

MiMo-VL Technical Report MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.252398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.252398Z digest=sha256:adb0d5b5bbfc641c24a7e2bba59a974998e5982845e20755fc1c807e16d42e9a

Observation 0557dbb1-cc66-48fa-902e-f0342f4081c9 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.256447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.256447Z digest=sha256:5e9d0e7116bd31383e5f12d4853baed7f138f6b5dd8445bfc3a57da671adbe36

Observation 6f58e01f-8ddc-42bf-84f5-462322cdc782 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.260139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.260139Z digest=sha256:f006f819d87aeacfa4be785af22b7afc1fd78ec7f1fd248fe272d25e4d1d3851

Observation 41847435-2c5a-495e-8d21-d61fea1e607a · outbound

This paper cites Demystifying CLIP Data.

MiMo-VL Technical Report Demystifying CLIP Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.263056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.263056Z digest=sha256:75e153c3ab1a84567b4cd0c46dbc256f64557b6a30e620818346194aa055b515

Observation 9c68946e-d713-4c77-a7f8-b444e217f206 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

MiMo-VL Technical Report Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.266387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.266387Z digest=sha256:469f792ef969a0a55b53bf9c474f1963c1c7ec230c8db48651b311258a591972

Observation 90ae787b-39e1-450b-a93d-7d962954c39d · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.851806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.269556Z digest=sha256:ddc93239c90aa8880a34d96e5972db03132249f008f58ae5dae83f45bca1b352

Observation d02eabf1-fe77-4d49-82fa-7cc796e5fc9d · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.841678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.272683Z digest=sha256:0d4c12a6466cbf2719d7ad4db00a0d4c8478b446c40689b9e90d4127f04e7485

Observation 03125271-f4c1-45c0-89b3-0ed29a95ad57 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.831752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.275700Z digest=sha256:58388e7b68ee8f3bad447236a3e3ca1bd40bf6aa6895a88884ff7d2b94e89e3c

Observation 3befae66-651f-4867-b1e4-fa7177b06cfd · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.821600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.278666Z digest=sha256:747222b43c0a55f907c1265f3808f048211c30a70d083786024e3ae3a3bf42a3

Observation feaed296-f60b-4e70-9790-c248f2ad0358 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

MiMo-VL Technical Report MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.281418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.281418Z digest=sha256:3bfc9d10bf2e2f8ddd064068ce14c8a80f9212c6b396a273d65f0ab5982ed9fe

Observation 399f54ee-1809-460e-a0d9-df04827b9876 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

MiMo-VL Technical Report LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.284762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.284762Z digest=sha256:b411ec4baa2fb3165864c6b89501fa1470a0f0752148dd1bd0c0f1b9afcc822a

Observation c3ef15d8-c794-4f89-ab02-c0dde79acb97 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

MiMo-VL Technical Report MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.288361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.288361Z digest=sha256:d667d36fce2739928fba6b22652d05ddb86b17fc024b100f8c61231a67c3bbb3

Observation 3d2026df-3505-4a46-96e4-998670e66d1d · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

MiMo-VL Technical Report MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.291602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.291602Z digest=sha256:736e516b468df8d84aa33888a73dc2d6856457b25ccce3cb51ff2ca6d0706818

Observation 67daf7b8-965b-439a-9e86-0d81d6023e52 · outbound

This paper cites Zheng, W.

MiMo-VL Technical Report Zheng, W

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.810708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.294892Z digest=sha256:d22d99ca6a0f32928ef747e2a5aad674d2a610008b74311a080512cc5cc8d922

Observation ecb9df86-496c-40c5-ad32-5a02543195af · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

MiMo-VL Technical Report Instruction-Following Evaluation for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.297922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.297922Z digest=sha256:c1a35a9a60ecf3ae490e7d39ca151108a57ea3a898ec6e6a22e5b388b1a148ec

Observation 7578a60e-6a0d-4f9f-9c08-77af19bdbbc3 · outbound

This paper cites Zitkovich, T.

MiMo-VL Technical Report Zitkovich, T

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.301233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.301233Z digest=sha256:7d01608edf7f3f5d041cbe49a9f95e1a68e3e2940520896d4d7fa3ef3bd91ba5

Observation 1b61006d-4166-4d49-ad64-4c018a689bad · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

MiMo-VL Technical Report DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.304184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.304184Z digest=sha256:68209f17c3d7f6a686af83fdf55e48b674c91603c482d327d9be6db8b2365f6d

Pith citing papers

Observation c28b2eff-3c30-4dc7-84c2-4f4f91a8c0bb · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback MiMo-VL Technical Report

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.014278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:2746b69c845f5e4e96c782692121513a47fdb33e94b9a8e72023fe3db7f77664

Observation 61343ccd-0658-4382-8a41-57d7baedba20 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MiMo-VL Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.993691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.993691Z digest=sha256:649f47adcb1d14aba6b6cde5175e33c39c3a8ba4b990f74f46687c13f40e2b7c

Observation b2c8541a-7675-4cd5-9fc8-2cb2f4bb202a · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report MiMo-VL Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.848642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.848642Z digest=sha256:4eecee101d70ac869834619429fdf7610929467e12cd78e62c16b682dbab00ce

Observation 643527ce-67f8-431a-85b6-f4d5b2f11e23 · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report MiMo-VL Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:04.898698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:04.898698Z digest=sha256:83bba88bc4945999db909404819ee0e5c27b08c4447cc4e616f932eaaa0c237b

Observation 46e0d4a7-b639-4306-8547-51f65d2169d2 · inbound

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study cites this paper.

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study MiMo-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:19:47.032597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:19:47.032597Z digest=sha256:42620edd8d78fbe0d4f867d1185d5846e7f57929ee79aa43656e2221d54fa5cd

Observation cf116d3b-3523-42a2-b6d4-71f203e48e34 · inbound

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents cites this paper.

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents MiMo-VL Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:32.939651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:25:32.939651Z digest=sha256:97decf7f3043a23ebcec02d8daeafc0339da3c0e22936a36fc783ce311154670

Observation 7b1b8f95-21da-41f6-abcb-9ef99a3cb01b · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback MiMo-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:34.638759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:34.638759Z digest=sha256:ca9bd57b0dd18faa4e0d061371eb0e7f1c957a9812361ad79189605721736d05

Observation 4d8ff12c-2c07-4a63-99f6-b547ba3544cc · inbound

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO cites this paper.

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO MiMo-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:40:09.085046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:40:09.085046Z digest=sha256:94e9161b5f0e1d14c2937c1c92db22ab551fe1583ef5d84dd619a1aeb34da6c3

Observation 71aee071-de03-4200-b863-2ceaaa125bea · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models MiMo-VL Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.109377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.109377Z digest=sha256:554f57870bc7869c6cdfc2e9137a07ed04d305f968607debed13fcf4ecaa03d4

Observation 9fd128e0-de73-428a-b81b-8ee18a9c45cc · inbound

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning cites this paper.

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning MiMo-VL Technical Report

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:16:53.749435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T23:16:13.165715Z digest=sha256:71a6b95417fc65256a9a037a4c8e6b432fea92bdee4f3ddcc3e1601a8e3e8f84

Observation 0b3a079a-698a-45a6-aa9f-76a8f375d4b1 · inbound

Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models cites this paper.

Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models MiMo-VL Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:15:27.996778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:15:27.996778Z digest=sha256:a6d4894057dfa8a6b176cb80bf29dd41556eb9f030b6dfd388febf5693a468cd

Observation 318485fe-dbb2-4117-bac5-cd3e05544081 · inbound

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning cites this paper.

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning MiMo-VL Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:21.112448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:21.112448Z digest=sha256:4ac56cf2ade0d10e869e10dfa56dc208b6d108e393ca7d42c83cf296ad08275b

Observation 1283b163-a25e-4989-99da-ea011bde4706 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MiMo-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.805838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.805838Z digest=sha256:d8ddf1b54951de8bfca4dbba30899c51fed9b85c6ce91116733601b7b827aa3f

Observation 83f477d4-a0ee-4e46-b838-1ef3770b3e50 · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report MiMo-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:32.039808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:32.039808Z digest=sha256:a8687a59a76047c8b42c7392c089a6a88dce935ed5cfcf955f5d7be8185225f4

Observation 66398f4f-8c4e-4ba3-b31d-33a74d5d8454 · inbound

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing cites this paper.

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing MiMo-VL Technical Report

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:06:50.111502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T20:04:00.773443Z digest=sha256:b50c9a74a53bdab2bc7b6eddc90c3570007c48f03c9812e28c3f3d54ac1db2c2

Observation 0c8d68e1-e4a3-49ab-ae44-cdf495f6c7a2 · inbound

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality cites this paper.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.883362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.883362Z digest=sha256:4eed99b30d497dffb74cb5a7621ba968a53cf91a5f3e9f76bc67b25413b48e33

Observation f71a09d5-6d54-4b7c-8b97-ee11e5773ca8 · inbound

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation cites this paper.

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.595790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:42.595790Z digest=sha256:ee4844e7c3c5f975407368a80af42eaf2e3b26acc3a8cabf11767817b4158ea3

Observation 297e34fa-3fdd-4af5-ade9-56901d741640 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models MiMo-VL Technical Report

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:57.154824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:57.154824Z digest=sha256:89be6596277c46fc8f23a12d13317ba7871faf0ed55f66182edb45cb95cb1b77

Observation b12fa6d3-9b00-4db2-95ce-68799fe4fb10 · inbound

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning cites this paper.

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning MiMo-VL Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T07:02:44.094748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:02:44.094748Z digest=sha256:f35329c4fff23c1c603745de29aa091abe51e59321a542ff059bb33d8ddba64f

Observation c6250233-8598-4bfb-9014-55b6ec3f5792 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report MiMo-VL Technical Report

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.687903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:c74eac24d0c6a63932e4833d384c65a263cded4c3e249c22caddb5e53b519af2

Observation c82b5776-93cb-4903-8d1a-afe3ddaf8a8c · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation MiMo-VL Technical Report

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.378830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:38:49.075457Z digest=sha256:23d9a58bbf7f4a5050f3ab078c7efca20b18280e91585138873d6b51b8903e91

Observation 73ec2059-50bf-408b-a4d9-affbcfca6975 · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation MiMo-VL Technical Report

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:24.167937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:14:24.167937Z digest=sha256:ad3f0099dd32620a7a1a2f8f2f4aceb29613e7db35f4f53b0b10235f361e67fc

Observation 4bd98a5d-fed8-4758-8c26-3f7e5d93553c · inbound

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? cites this paper.

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? MiMo-VL Technical Report

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:40:12.591587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:36:28.844714Z digest=sha256:6c89504c85d59f6cef08e312e572103b3a56579cf78e5c1bad63813d4266fb51

Observation 8628686b-b34e-440d-8232-120cf27510e9 · inbound

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? cites this paper.

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? MiMo-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:31:21.274478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:31:21.274478Z digest=sha256:7ab0b6e9dfca3663302ae37a5782df75fa3de9a1c08a9ab51b95d2a309d1b669

Observation 9f537819-a78b-4e4e-8494-e6c407fc8e08 · inbound

Learning Self-Correction in Vision-Language Models via Rollout Augmentation cites this paper.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation MiMo-VL Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.930637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.930637Z digest=sha256:ad6b981fe7746419809b4b57507fa18039faa67b8b8b768161cfa18c4511afa0

Observation 9a820de4-a910-4a8d-a76c-6440165bd6fd · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.621170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.621170Z digest=sha256:8c0fd5ead7fee98170b7ea00a1f66ac9d31cb5ba38897cdc705684ab2dc01945

Observation 54be8578-9b88-423a-8e54-6368f2bb3019 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards MiMo-VL Technical Report

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:01.057349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:cf7af8e5fad8c1057be7dab0f6aa846ff75654c8643e177e5ab0655dfd366306

Observation 2ad5634b-2940-4275-8fe7-d9223b08657b · inbound

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations cites this paper.

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations MiMo-VL Technical Report

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.855125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:23:45.814042Z digest=sha256:05e8303a5e253d64a34fcde940202d54625d0208dda1a83e0f67c8931a093dce

Observation 93ddcea5-acf8-4ffb-8237-018a3f76e2a3 · inbound

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models cites this paper.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MiMo-VL Technical Report

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.615360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:1cc11bef8399a16b3ab41788bd218734699a6c812fbcf5eee696d9f33ed4af91

Observation ccbc11e3-b927-4d3a-8df1-943a753f759e · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model MiMo-VL Technical Report

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.261247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:d0580bd7fc02175984e6e7504470fe810ccb95115fe8c83e3e05516d3274ddf0

Observation a0dd9a59-629e-4a98-aa0a-89287382ae77 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MiMo-VL Technical Report

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:43.902486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:4aeaffb7544751075883d3ee0ac32bc37ba9d77e151804ead9a7132a2cea620a

Observation 77cabe27-51d5-40c4-aa38-b3a2361805c2 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MiMo-VL Technical Report

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.700497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:e5245f01647a08543c61d881cadfbb0dc3c3556e6ee5b9a217486ad68e24a1c0

Observation c4354bf9-2a8f-4871-91f3-eb6334a7b9e1 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning MiMo-VL Technical Report

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:55:44.279701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:20b163711b43d0871ce13792eb6b928800aff05d1f128d6a15177dfcf3d331c1

Observation 0f5b740a-da5a-4aff-b708-a435da384cba · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos MiMo-VL Technical Report

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:55.990591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:eb36c9014a29c2395bbbbaa07f92adae5cb9ce00f62e84e742c0135eb05f61c4

Observation c894e30b-280e-44aa-b709-a1bfb5edb5a9 · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models MiMo-VL Technical Report

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:45:58.468198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:f28dcdf07f2ddf778dd63cefb36384979f179e87d7fddd368813332b8d7316be

Observation 893cbb9f-1030-4cab-b4cc-941550d89e86 · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:37:03.526884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:35:09.568623Z digest=sha256:6e6416b4c51c9cd7a4aba6b78c7fce56436316289b9001d2b5bdfb910eb199b9

Observation 3ac97830-2507-4a0b-98ad-ce6f04e24b6e · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:18:00.124492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:10:49.059686Z digest=sha256:6f1d829420022cfe7706f6eb170735dfdfef733fe5e30da2e5d6c4aa22732293

Observation 0569c09f-ce8d-4b67-80c0-7a74e654719c · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.597384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T16:58:55.817334Z digest=sha256:bab8e2c7a50d5780b26ef348a0ef51b7043d288f95bffce89874cdb8b8fab4c7

Observation 031d66aa-8f0a-4a00-b416-864604493087 · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:46.445675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:41:13.164195Z digest=sha256:81eda5a1226f5b355075a1785b66ace5df6e8e48195a415f7d94ef870c150262

Observation 0334a5b5-72d2-44bf-8c11-5bb24cc56331 · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding MiMo-VL Technical Report

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.412313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:2ec18ed5877ebf3c440da65c4e34fece797483b03c798283d84dccc5dfaeb51f

Observation 9e181396-0d43-4b55-809d-c86d586c3484 · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation MiMo-VL Technical Report

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.489809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T10:58:01.621488Z digest=sha256:cb49a726f91eaccbcc5c805dd1c238a6d2aa801d59d4481c1805eba7a05e2177

Observation e9a0c321-523d-4464-b0a9-c3eeeeffb87a · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation MiMo-VL Technical Report

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.506207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T18:28:21.605646Z digest=sha256:acc4fc8646131f931ca023d54e8521b070b0b04b965cfd09aec7036bf4c0ac58

Observation f91bb6c9-f935-4c34-ab58-92281d03fbf3 · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues MiMo-VL Technical Report

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:14:42.401670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:766a6196b461d500c69036d31e19e29809988540b21a508a0360eeff32caff5b

Observation 83e8f609-e1d0-4b50-9f14-5bc4fcd3faf0 · inbound

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding cites this paper.

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding MiMo-VL Technical Report

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:06:15.393088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-22T08:04:55.680668Z digest=sha256:1cf6e13f70a92cb30b94b8d46e6057dbc3e60f6bf329b461297874ec07557b27

Observation 7fa44757-a829-4dcb-9436-49bcdd706b9d · inbound

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis cites this paper.

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis MiMo-VL Technical Report

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:54:43.935849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:51:47.455642Z digest=sha256:83ddee269c5c748ebbb3e87b20cebd06a7e8ffbd5adbc6e6b4117e09250fd859

Observation a5a691c8-6b18-4d0a-91c3-a657e0f751e0 · inbound

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker cites this paper.

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker MiMo-VL Technical Report

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:02.166045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T23:15:02.455554Z digest=sha256:b5a4704ca9c147306872eca004e83eabc0a24c3cf5d71852c6f97d4ac081032a

Observation a6ec13b3-b844-40b6-8ed8-9bcf9aa26393 · inbound

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning cites this paper.

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning MiMo-VL Technical Report

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:42:46.667492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:38:30.213948Z digest=sha256:ca9985146f5ac0ac063bb61b90398389607702eaa1f2c8ad228558fe7ec60195

Observation a3b539c2-eae4-444f-ae51-e861be05b60f · inbound

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL cites this paper.

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL MiMo-VL Technical Report

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.971431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T14:55:07.045625Z digest=sha256:a35b353d69d0f51bee1d818ad8fa29cf13504b041091b4dcd5dea9273424c6f6

Observation af881701-d82d-4273-839e-3ef5afd8e92e · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding MiMo-VL Technical Report

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:27.944526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:83a81aac9b5269dcacc9fbfa99184cca2c7807a8f08ff617f8fd777e97129617

Observation 9ba9fe9f-ca95-473a-a74f-74f5ed8b7c7e · inbound

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues cites this paper.

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues MiMo-VL Technical Report

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:47.984436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T06:04:28.939248Z digest=sha256:b2294443c53fef0448e76bcecae570960f29bfcac91a90d0d91916207786f68e

Observation 229438a3-c704-4337-85dd-68bfdd291a7a · inbound

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding cites this paper.

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding MiMo-VL Technical Report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:24.811296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:40:56.851297Z digest=sha256:1078e83feafaf7f6f80927d4b6eaa376f523f4c93684850922fb2b9213f6b260

Observation f312811b-cd21-4be4-9a69-dd19e9c8f32e · inbound

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding cites this paper.

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding MiMo-VL Technical Report

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.033020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T10:04:29.739632Z digest=sha256:678c0068f7e2b8d3bb61f7f311708556600daa7363efadbe8235789c72ade0a6

Observation e8f54eff-88c8-4653-95b8-f558789e19cf · inbound

AIR: Adaptive Interleaved Reasoning with Code in MLLMs cites this paper.

AIR: Adaptive Interleaved Reasoning with Code in MLLMs MiMo-VL Technical Report

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.983706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T09:06:38.001604Z digest=sha256:d9c278fce9008f59e8d243fade44b1acc9d9c1a3d1d64ad57d004495875ff6ff

Observation 6a3f5543-80b6-4747-bf82-9c2e3d3ef69f · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning MiMo-VL Technical Report

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.283757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:60f51d846a9708c9a54cbe39d54714cbf18b0f3cdf37ae9d5e928bca0191af9b

Observation 991bc27b-06fc-4768-8e66-36a67bbedc2e · inbound

Aloe-Vision: Robust Vision-Language Models for Healthcare cites this paper.

Aloe-Vision: Robust Vision-Language Models for Healthcare MiMo-VL Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:57.944069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T02:02:47.472868Z digest=sha256:6d1c7cdccc0224aed9cc9243cc40903aae88fb5085af4be82aad5a0d8c828efc

Observation 36501310-a408-4f0d-89a1-e547603e6e68 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models MiMo-VL Technical Report

Reference 277

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.594634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:2fdb7c0fdfbab095b25bdf7f98df9c9b2741fab49dfd74796eb165af84c6e1b5

Observation 485ee3cd-1f75-48bb-ade0-16c332cb0008 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models MiMo-VL Technical Report

Reference 277

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.078871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:698f63e28b65ea628c0ac7bf81217a71fcdc36d9612a75d402b341561781ed5b

Observation e3295a0c-7ae7-45dd-8c0d-f85951fb995d · inbound

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension cites this paper.

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension MiMo-VL Technical Report

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:48:35.007245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T15:41:03.548799Z digest=sha256:ee03c698cf7a0fbdbfd5df3e6665859dfe89aec62f8f606222ae0b51399bc6c2

Observation 45c10caa-2eb0-4a62-b485-f74be4e92864 · inbound

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools cites this paper.

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools MiMo-VL Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T07:02:55.388823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:02:55.388823Z digest=sha256:75afd4267e5e5afd3fdadce703051a9bb43d4f56c355fbc4e7bed74c30ef4dba

Observation 1fb177ae-653c-4280-b4d2-86794a0e404f · inbound

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence cites this paper.

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence MiMo-VL Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:30.687572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:30.687572Z digest=sha256:065bcd1455e9246679208ed7b74360adc312bae329a5b19c84af4d2b0c99cc6b

Observation 8c9a65e3-14b0-4b3d-bf4e-54edb7612293 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs MiMo-VL Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.011960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.011960Z digest=sha256:0b6fa7c2b4689bf352f2253695ec8e1276ac6ac59c53ff31e6e1df1cd66076dc

Observation bc403541-2b97-4b33-9b74-c58a6aa4d875 · inbound

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement cites this paper.

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement MiMo-VL Technical Report

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:43.966698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:43.966698Z digest=sha256:9c29b118d06ff86b6278735a755019c2ffd66016f53b449ae41e49da7950a900

Observation 38a6600d-e2b8-4c86-bd0e-470aeb0b9fee · inbound

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis cites this paper.

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis MiMo-VL Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-30T12:15:47.193300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:15:47.193300Z digest=sha256:b04927f51e9aa0b087ccb6fbccd9a992d7426ec87cac5bbe772588bdd7b6949c

Observation 72382549-ea31-40eb-a9bf-70cb567513b8 · inbound

NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model cites this paper.

NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model MiMo-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T09:49:33.571423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:49:33.571423Z digest=sha256:ac0687faac6430d4c9c162a6a2d63e44629863b97811314a6ed056b8b393dbfd

Observation 958d1bf7-dbaf-43b6-8f45-47cd133a3c08 · inbound

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research cites this paper.

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research MiMo-VL Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:16.891205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:16.891205Z digest=sha256:39550d2d5b6a642eb90ba59900113e7fc68e1cfbf0cc958f20438c280d561766

Observation aaca9b6d-975b-4fba-a7c6-4d74fcfd061d · inbound

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes cites this paper.

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes MiMo-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T01:14:37.781329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:14:37.781329Z digest=sha256:0b788cc51c1a6ccba0ac5c3717f92ae3733d70bc7fb3442da82c4adb643330ab

Observation a7a12be9-d250-4a05-8de4-ccfca55f846b · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning MiMo-VL Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.095639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.095639Z digest=sha256:63af460d890df4893bbfee28a195c732c91732508a06b3320a430d8047f71a6f

Observation 1c0ff46f-77e5-48d1-af0e-05b385705b8f · inbound

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding cites this paper.

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding MiMo-VL Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T15:43:56.781142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:43:56.781142Z digest=sha256:33b61b95cdba5c8a2a4c8e500598635a6bfb9bc5620d5c227c852130bfc14dd8

Observation d5d883eb-ec4d-4405-ad37-cf7b0f03d325 · inbound

OPD-V: Visual On-Policy Self-Distillation with Modality Balance cites this paper.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance MiMo-VL Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:15.205510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:40:15.205510Z digest=sha256:8d706a70e1f3f9ae14ae60b85eea16b61e115b32fd36b64a1e569d2807d22d0e

Observation ebcee3e8-f82b-45e2-ad3c-d79c23b1e969 · inbound

OPD-V: Visual On-Policy Self-Distillation with Modality Balance cites this paper.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance MiMo-VL Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.830087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.830087Z digest=sha256:120ea6c6b364c69eca17502e084ced7b2fec597e670b8e5bbe5027e9fc0e89e5

Observation 7a18e0cb-2629-4751-b8c9-9c77aa6e7cc3 · inbound

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning cites this paper.

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning MiMo-VL Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:59:13.756188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:59:13.756188Z digest=sha256:e8d378d6cec29147e9dc1b663d1f14abe6a73eb8fc2681fe152b9ff70723ab5f