Pith. sign in

Paper Citation Record · LEDGER

MiMo-VL Technical Report

As of 7 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 69 inbound Pith citation observations for arXiv:2506.03569.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03569 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:14.304184Z

measured 144 of 144 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 69 of 69 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:22:55.993691Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:57.282206Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9aa704c6-9745-4a83-9db6-73fd870606d3 · outbound

This paper cites Alayrac, J.

MiMo-VL Technical Report Alayrac, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.116017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.025617Z digest=sha256:4bb85e0568bb23d8693ea612040c4fc9d388cb9abae5ac62ada8f2c83c23aae3

Observation a2826d57-f206-4ba1-be60-983d963a0685 · outbound

This paper cites Qwen2.5-VL Technical Report.

MiMo-VL Technical Report Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.034077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.034077Z digest=sha256:209c8baaa45cea5ad127826969d84498fe42a759c6412b5b1c48d2bb85207df0

Observation 50a1dfbd-b5e8-4c7c-a3a4-32c2cb1cd12c · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

MiMo-VL Technical Report $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.038069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.038069Z digest=sha256:1f30d5a036a0c14ee11333a2a24fc835f684521ccc2efb6f1484e45445a0dc9c

Observation 670331bf-23e3-4d08-8d1e-a9d3fb28f477 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.105797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.042192Z digest=sha256:12bac9b8e406648cb409a52a8bba11d7faab62f60802e803fd48b4a173316511

Observation 3ef05f39-56ca-4db7-97e5-7c698381868f · outbound

This paper cites WebSRC: A Dataset for Web-Based Structural Reading Comprehension.

MiMo-VL Technical Report WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.045891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.045891Z digest=sha256:2601054b15fa16af5c1bc69c7c699971616e29147e47d19b544112cf26e9ccad

Observation 106c8006-b943-42f6-89be-79b3fc2895ba · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

MiMo-VL Technical Report AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.050518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.050518Z digest=sha256:20844ea8d11e2dd25e80d962ff70941d87a4342a51686fb94ae8778af326e7dd

Observation 6488aff5-0583-4f75-86e2-4a7535c09af9 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

MiMo-VL Technical Report SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.055373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.055373Z digest=sha256:c5e1e6b7aa2f33bf315de61f7a6ee9df824ef4686258a601c7b920585e55b866

Observation 6ff283ef-5fb7-446c-934f-e085a2eb0250 · outbound

This paper cites VisionArena: 230K Real World User-VLM Conversations with Preference Labels.

MiMo-VL Technical Report VisionArena: 230K Real World User-VLM Conversations with Preference Labels

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:04:14.747082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.059734Z digest=sha256:ae230de5521340cc8524adb826b7f3f276276fa965d0c6c240c6cd26c7965dda

Observation d1f6d652-58c4-4a40-8837-16795209b209 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.063690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.063690Z digest=sha256:c04b0105ccbe880464bce426d4f1744b138031b80de15498f828b51e98aa90f2

Observation 4c1e6125-b66b-4986-9ca1-38897496c9e4 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

MiMo-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.067321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.067321Z digest=sha256:b25ce23101854b478e0128f1c8130d442b309735e6b9f8537b2008ea513dc536

Observation 8798a9b4-020a-40fc-8455-e5210490734d · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

MiMo-VL Technical Report SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.070645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.070645Z digest=sha256:d6947519740327c4bf0c8a0dc99c20d5cc1a34332551a044427358d3c05d7a26

Observation 298525a3-dc3f-44c8-bb97-016218fff5ec · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 13

Resolution
verified exact
doi, observed 2026-08-07T11:04:14.336562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.074224Z digest=sha256:af0825cdfd791b3995b170fc101efc0542f7c189fcf8f08383df58507433e2e8

Observation dc824662-9735-4f92-aa61-5b0ab13b7362 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MiMo-VL Technical Report Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.077563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.077563Z digest=sha256:796d55a8e7bf302623f4674bbef16eb1048ce6a3cc1de9c6919a64c57f0a80f8

Observation c985dc53-33c1-440a-945b-883a04ae9c81 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.090052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.081196Z digest=sha256:30a29e45881671ecd5f82aef39cf3a2c19e204f7f4a2847948685431e3bea086

Observation a4c3bb18-9b57-45d0-bfea-258c3bcfb8ac · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.085568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.085568Z digest=sha256:b7b183b88bf9d386f09baed935244b5ac769948f1c9d48c2740eac24be2f4965

Observation 747b7261-a0d9-4748-8ab5-e02023921a45 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

MiMo-VL Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.088974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.088974Z digest=sha256:6f08d9c8c987506459b12ad492c1c28897ac07abd7982b1a22ed1678559bcb55

Observation 987da426-a6e2-4e58-89e8-0e770f4d570b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MiMo-VL Technical Report Measuring Mathematical Problem Solving With the MATH Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.092481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.092481Z digest=sha256:a47b6462ed7d471518efb0c10a020c7e97adb97709b5badfb9445e5a86877aee

Observation d211c42b-33cd-43bf-a248-ebabe66b7dc9 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

MiMo-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.095997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.095997Z digest=sha256:a821a967373ce7664c119ba92f198c7c8c476e00bd02a20fdf1e874e43a1f3ba

Observation 483754e1-be5e-4ab1-bffb-67493c08e6f3 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

MiMo-VL Technical Report MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.099478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.099478Z digest=sha256:90865136f7483a09aa6db12ba5f1ac0c3e1b0a59c7ef8a53d8d5bf1151674e94

Observation 4c2bb9ce-966a-4916-bb14-8d66bfb6ff6c · outbound

This paper cites Karamcheti, S.

MiMo-VL Technical Report Karamcheti, S

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.073290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.103006Z digest=sha256:581e69ae7d1bca1ba5b43bb64e4643feb0c9418b66d2890bc4603dd4e19ad6d9

Observation b54d876d-a504-455c-a503-e2fdd9f8b944 · outbound

This paper cites Kazemzadeh, V.

MiMo-VL Technical Report Kazemzadeh, V

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.062551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.106351Z digest=sha256:6e4a965397cf3fff676aa63844e70a6c7102d32e2faf13574ccc2cc75ecce861

Observation 38e4f808-d3be-46c9-936a-82138f84335e · outbound

This paper cites Kembhavi, M.

MiMo-VL Technical Report Kembhavi, M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.052054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.109684Z digest=sha256:38a85e87a28ffb17947ac5dd5443eff2e8137c8fc673bff9b32bfed3d2479093

Observation da410722-19b9-450b-b315-c59a0296f470 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

MiMo-VL Technical Report Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.112913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.112913Z digest=sha256:b43010a0c71f5601e275a88bc18bbbbe2fd649472f37c0396b5df75cd26ce04a

Observation f24a4638-6649-4dad-9272-65e048d80726 · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

MiMo-VL Technical Report SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.116722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.116722Z digest=sha256:f1359215ddd5999253da7337eae8598f90d2ca6e75bc232ac2595f9eb1d35393

Observation 05a081a7-53ea-40d3-a02d-6246af5351d4 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

MiMo-VL Technical Report ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.120139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.120139Z digest=sha256:ed359f354bce0fd9d0ca0039e6901e66738be97b5b16dbdaec372837a86c5c38

Observation 350c9168-2d72-4d6f-92d7-aba2d8543642 · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

MiMo-VL Technical Report VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.123685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.123685Z digest=sha256:b2d7de8764071f21ccf0501524119796bc781f88c627a12e50eca5727a59e923

Observation 77e454d1-eab0-42c5-9876-16cbc6fc0224 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.040120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.127539Z digest=sha256:4a3593a06900291a383a488247990bd55a2580f43b1fa9c756f03ef4bb571e96

Observation db0c7814-3e4a-4b6f-86b6-8e5a0d247c6d · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

MiMo-VL Technical Report VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.131074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.131074Z digest=sha256:fc6706759f79600f5c8ba8087e8527dbdf6ae3722c9c9b3326664f2e6f98d925

Observation c8b776c9-bfe4-4024-868a-d8d1b4f460a9 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.030409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.135029Z digest=sha256:d71bf4fd71dfaff4cb9146dff81dbf1cc812f763f3d25ae6471648d0a63167a7

Observation 50380670-97b8-45ca-b403-70f4ae033c09 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:15.020268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.138906Z digest=sha256:0364192d4d097dba7e001138c0bf4eb107204230f338f01317e39c959a723cab

Observation 6c7eed2d-cf0e-41f6-8551-7053f830d656 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MiMo-VL Technical Report MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.142694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.142694Z digest=sha256:21b62304227537dc01edc074a50b909fde2446f78089729d22d29c6acccc47c6

Observation 4ef01348-6663-492a-9c0e-ae9a7d88f4e0 · outbound

This paper cites American invitational mathematics examination - aime.

MiMo-VL Technical Report American invitational mathematics examination - aime

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:15.009244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.148173Z digest=sha256:50e6bc04f08fb4d0b8e948f9d60e48d26256bbdd0896eef68fc27c374d481e27

Observation 5ce0a271-881f-47b4-8ed1-e218205e07d2 · outbound

This paper cites American invitational mathematics examination - aime.

MiMo-VL Technical Report American invitational mathematics examination - aime

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.997894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.151937Z digest=sha256:c949c4d0c68f508b7b7ae09767487836433aa11f1f57dfc3ca875754fcc3e12f

Observation 05505eba-d9a1-42e4-8397-5fd8b2c0c4f1 · outbound

This paper cites Mangalam, R.

MiMo-VL Technical Report Mangalam, R

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.988056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.156916Z digest=sha256:261647b978e1d35040771c73d20542cfd7dfc48566a71bb73a2a501eb420b332

Observation 672eaaeb-cfdd-4f7b-8ac2-88fcf873d575 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MiMo-VL Technical Report ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.160624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.160624Z digest=sha256:0187cabbde8b9b9b5e943a3017c06b9a42dd9e143ed35ed0ad26c62a569c764f

Observation 7409a732-62cd-4b2c-ad5f-db677b1d8397 · outbound

This paper cites Mathew, D.

MiMo-VL Technical Report Mathew, D

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.978055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.164676Z digest=sha256:eaf94664f8c1d78e4ec87b02b517ebca629ac8784f63cce089bb47a136a36e5f

Observation 0d41bc7b-b4d2-4d45-9f17-79c5146d37ac · outbound

This paper cites Mathew, V.

MiMo-VL Technical Report Mathew, V

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.967115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.168315Z digest=sha256:ed76b14212e3f5fe8bd1c6fd2ba057129594e383b1d94d953cd094e993bd2ff5

Observation d0b5c300-4c4b-4f15-b85b-fa571334f0b2 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.957073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.174943Z digest=sha256:1a622cbfda335bf0b865869b14b4b929f5502a0e75e06947b8c7afbfd437192e

Observation d9e6d50f-7a39-4fd0-a161-ca5687bd38fa · outbound

This paper cites Computer-using agent: Introducing a universal interface for ai to interact with the digital world.

MiMo-VL Technical Report Computer-using agent: Introducing a universal interface for ai to interact with the digital world

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.179094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.179094Z digest=sha256:1982fd5613f423227399f7faeefce17bd6d98c325ecd51dda727dc11812fc588

Observation 38547bc7-66bb-49a5-973d-4edba5f007c4 · outbound

This paper cites Ouyang, J.

MiMo-VL Technical Report Ouyang, J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.941293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.182835Z digest=sha256:4d0d115927a8165e70e328df51d19f25fc2d5188fd7f2eb13fae1701c93c56d4

Observation 393c8537-f874-4d32-8226-f091b85810bc · outbound

This paper cites Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models.

MiMo-VL Technical Report Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.186518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.186518Z digest=sha256:cc3763d536468137bc19a3cd355836ce07a53719f067a5252d6a97cef7eb6029

Observation 1dfcd675-3d31-4dbf-b1b1-035ce52f7aa0 · outbound

This paper cites Paiss, A.

MiMo-VL Technical Report Paiss, A

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.930841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.190120Z digest=sha256:286d4ddec2afcb75f74c6651acf349f5af7fd0c45b5277dd990207343b6509f7

Observation f93ff15b-64ec-4967-b269-e728eda1e8f6 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

MiMo-VL Technical Report We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.193506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.193506Z digest=sha256:6d012a44babcdabf2830495e1281ddf8c4ae023a5820868585e6d0e48b739336

Observation 20a32917-7a98-4a22-8e4f-fd16a6b7521b · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

MiMo-VL Technical Report UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.200543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.200543Z digest=sha256:4014c0f1aa17ab5cffb128823b38fdd0e58f219488a196a7a7a63ef5bb61666e

Observation 4a7c6114-3b17-4b08-a388-ef2a0bdc0a10 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

MiMo-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.204341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.204341Z digest=sha256:8f0043eca8dff29ec125a566c81ab1faa2f5b944673e95c03ca50bb91d5ee19e

Observation 1a87f416-d202-47dd-99c8-72f7bfeed1cb · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.208120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.208120Z digest=sha256:b3fa405fffd4c348b735e3ca2eb22505a2d9501fb3f1b8de07b0276540275ee6

Observation 194075af-6572-413f-9fb6-d5da771e53de · outbound

This paper cites Rezatofighi, N.

MiMo-VL Technical Report Rezatofighi, N

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.912787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.211431Z digest=sha256:77c48d5bf2fc0dae4a4a21d8f404d872f25f8fc256ed6f3c0029a6997c6d4631

Observation f7c35073-bd3c-4415-9f5e-2b0e483d572b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MiMo-VL Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.214574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.214574Z digest=sha256:7c4739859f0e397cbd52cb7f86b93ac4e2f39846d0d8febb0cf6798697d00179

Observation d1886f12-039c-49d5-9267-da3c22794b7c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

MiMo-VL Technical Report HybridFlow: A Flexible and Efficient RLHF Framework

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.217839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.217839Z digest=sha256:1d1f5892bf2af9747e2e5ee6ffe2fe693b14bd0a6540b5404bed858f6b69499f

Observation 8d03c616-b330-429d-bf9b-6b6a85f18dd8 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MiMo-VL Technical Report Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.221120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.221120Z digest=sha256:6094cd2e0fded96ccaa58dbb64e63e7407d334ed1d434ca4dce70021d8a4b523

Observation a3fcf5b8-5082-45d0-836b-ec5d85db990a · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.902472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.225004Z digest=sha256:87dba0bcdf9ba569d59d332e73f75f12e46453bcdc7a2ff6a7ababe1dad5343c

Observation a307c3cd-758c-45b7-a4c2-7b57ef8323ba · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.891258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.228284Z digest=sha256:f8a0a674c202e166f6729c74f9c1516736e75ed4f4cd3af7b78d6e63df138d5f

Observation f5df548b-aabe-43e9-9ea2-31398e8463c5 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.879951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.231882Z digest=sha256:1435f2316ebba60929848e76b28f200895e05f170c422afc534dcdbfe1c74c95

Observation 12058bba-b8d3-4e0e-bbc5-eab30b9b953e · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

MiMo-VL Technical Report Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.234976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.234976Z digest=sha256:3ef5fc1317ac6a62ba14ae06b3eec4b90de7d9afb104af357bfd495a1ed9a5a5

Observation e708e028-114c-4a3e-b3dd-ed2805a5e282 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.869525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.238436Z digest=sha256:c351fc971ee46793152d5e6c4398207172561939a41d40b91cce8eef1fb3ce40

Observation 7599ee8d-24d9-4414-ba22-ad3aa2d74bf5 · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

MiMo-VL Technical Report V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.241662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.241662Z digest=sha256:a46973b9e02b09326e23120ededeb1a3f76b054705d78380aa6e52aa4b3cbe0a

Observation c9641ad7-43ab-4fd8-919f-7e04dbdcd212 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

MiMo-VL Technical Report OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.245410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.245410Z digest=sha256:5f96f54d637b08c583c61042679df8e1180db87152601f8b0a20ee63e05cfd8a

Observation 80a3932f-3008-446c-b213-fbec0feedcc5 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

MiMo-VL Technical Report LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.248887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.248887Z digest=sha256:7046c418602b1065475403288022b6775288b4151bd2d5d481b9ed2ea8c212df

Observation 64865b0b-3a69-4a7f-9f3a-9ebbb71b8cf6 · outbound

This paper cites MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining.

MiMo-VL Technical Report MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.252398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.252398Z digest=sha256:7a8408fd3d37508790e5ae60df0054a42fba2d28d54ec6a392bdbcb1b327690f

Observation 0557dbb1-cc66-48fa-902e-f0342f4081c9 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.256447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.256447Z digest=sha256:53f7e1bdd6a132eb2cec52b6b481f525cafa3f81cdd135f7612e96a9fe1928ef

Observation 6f58e01f-8ddc-42bf-84f5-462322cdc782 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.260139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.260139Z digest=sha256:e477b30b1fe8a111c12ca2fac3adfb3cf0e1fa15745a7cb036e04a5140b9e4a8

Observation 41847435-2c5a-495e-8d21-d61fea1e607a · outbound

This paper cites Demystifying CLIP Data.

MiMo-VL Technical Report Demystifying CLIP Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.263056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.263056Z digest=sha256:b6bef98f599d3df506005b2e2c6f5957bdbd3471ed8b549172eda8b0df8abec6

Observation 9c68946e-d713-4c77-a7f8-b444e217f206 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

MiMo-VL Technical Report Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.266387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.266387Z digest=sha256:2265e02d806b60eb5338a41402857a8712ef108f4dd99f0f3910b615c233814d

Observation 90ae787b-39e1-450b-a93d-7d962954c39d · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.851806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.269556Z digest=sha256:75096be70900074a456454517bc1081329be4e9d36a05c565cba8685c139c23b

Observation d02eabf1-fe77-4d49-82fa-7cc796e5fc9d · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.841678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.272683Z digest=sha256:69e894cfb36335679dfaf58ae438b181feb97af8e4beba7c51884fcafc59f57d

Observation 03125271-f4c1-45c0-89b3-0ed29a95ad57 · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.831752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.275700Z digest=sha256:5a3b70874cbeeee23da4ba9e5980b9da02ccbd3fb29c76e305db883ff9b70623

Observation 3befae66-651f-4867-b1e4-fa7177b06cfd · outbound

This paper cites an unresolved cited work.

MiMo-VL Technical Report Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:04:14.821600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.278666Z digest=sha256:8ce7497db5df907c02b0181f4fdac3454b1104bc4ffc9dcc06a5f5fb81e36555

Observation feaed296-f60b-4e70-9790-c248f2ad0358 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

MiMo-VL Technical Report MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.281418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.281418Z digest=sha256:691909c42ca6c3579a3f1332dc1f154e0c1de7223152671ca22eb1b9f65d05b7

Observation 399f54ee-1809-460e-a0d9-df04827b9876 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

MiMo-VL Technical Report LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.284762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.284762Z digest=sha256:f6e3414a4e6310e8bb7274455d9f8fc14a06f9ae6becf200e506fb73a3c9cd60

Observation c3ef15d8-c794-4f89-ab02-c0dde79acb97 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

MiMo-VL Technical Report MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.288361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.288361Z digest=sha256:56741dc92067a00d5f6c10389818f927162b92fc00ebe3b392019d3675957bb3

Observation 3d2026df-3505-4a46-96e4-998670e66d1d · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

MiMo-VL Technical Report MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.291602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.291602Z digest=sha256:d15045eea2ec2405811cfe48150644991893adde0dfc56acc18f6852c4149e75

Observation 67daf7b8-965b-439a-9e86-0d81d6023e52 · outbound

This paper cites Zheng, W.

MiMo-VL Technical Report Zheng, W

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:14.810708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:04:14.294892Z digest=sha256:63e362fab8fa2d0139ff6eef05eee43c32ade77217611013597208a89c00ac4b

Observation ecb9df86-496c-40c5-ad32-5a02543195af · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

MiMo-VL Technical Report Instruction-Following Evaluation for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.297922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.297922Z digest=sha256:77e42113c79bd3cb0034e41e0be1173787bc11a3ee5877a789e3bffeb09eae9b

Observation 7578a60e-6a0d-4f9f-9c08-77af19bdbbc3 · outbound

This paper cites Zitkovich, T.

MiMo-VL Technical Report Zitkovich, T

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.301233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.301233Z digest=sha256:2afd51f1a0c75ca3884d5036f6224f0e1c93eedca84d98e51053105702762c35

Observation 1b61006d-4166-4d49-ad64-4c018a689bad · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

MiMo-VL Technical Report DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.304184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.304184Z digest=sha256:84f4a0b27107dbb0ac615935dd34ae30fe711512f08059b5dc743feeac884083

Pith citing papers

Observation c28b2eff-3c30-4dc7-84c2-4f4f91a8c0bb · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback MiMo-VL Technical Report

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.014278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:392997202fc1742881a926cffda7d2c19cb3c6b9cb3280db3adf5c8288c90ce2

Observation 61343ccd-0658-4382-8a41-57d7baedba20 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MiMo-VL Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.993691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.993691Z digest=sha256:42d8a698cf014932d8283acfcc7273e6c026bed7f5a01bedff7f82a9a6238a03

Observation b2c8541a-7675-4cd5-9fc8-2cb2f4bb202a · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report MiMo-VL Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:09.848642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:09.848642Z digest=sha256:d18a722f9f3cea1a0e042d8b576c4e10f5a03658b7fc0be90633786a4a6b8e18

Observation 643527ce-67f8-431a-85b6-f4d5b2f11e23 · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report MiMo-VL Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:04.898698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:04.898698Z digest=sha256:d3eba6b3b4cfdefb2df6c2b3e9fe70b3d059863d4a97bcb4e01368fb9ad6d8a3

Observation 46e0d4a7-b639-4306-8547-51f65d2169d2 · inbound

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study cites this paper.

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study MiMo-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:19:47.032597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:19:47.032597Z digest=sha256:54bcbaaf8fc1f3a847941ccbf8866614bca37adb69bf55fe06252b2d149b58be

Observation cf116d3b-3523-42a2-b6d4-71f203e48e34 · inbound

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents cites this paper.

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents MiMo-VL Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:32.939651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:25:32.939651Z digest=sha256:d59be3dfa775a246a7a1619a239402c0a01980ac172e2215e669983160d657da

Observation 7b1b8f95-21da-41f6-abcb-9ef99a3cb01b · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback MiMo-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:34.638759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:34.638759Z digest=sha256:1459bf12536754b3261fd0571e18652047e2a0e2599adb843d5852d423240413

Observation 4d8ff12c-2c07-4a63-99f6-b547ba3544cc · inbound

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO cites this paper.

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO MiMo-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:40:09.085046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:40:09.085046Z digest=sha256:568452dcc6f836ac7f7a3e83b641d36cea7916cc2f972cf19ec1521282462cec

Observation 71aee071-de03-4200-b863-2ceaaa125bea · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models MiMo-VL Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.109377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.109377Z digest=sha256:6b096e87c1f6ff86f14a3bb6b56b35080c4983c0608bc247f625a08847344818

Observation 9fd128e0-de73-428a-b81b-8ee18a9c45cc · inbound

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning cites this paper.

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning MiMo-VL Technical Report

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:16:53.749435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:16:13.165715Z digest=sha256:84bd4bf4050343fca473feb323da1640c3ac9f2f47499ec93f16c89c08f81aaa

Observation 0b3a079a-698a-45a6-aa9f-76a8f375d4b1 · inbound

Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models cites this paper.

Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models MiMo-VL Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:15:27.996778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:15:27.996778Z digest=sha256:0e5f092b5f9ba60e8fe5a9ad9a62c9ff514a138ab898da6b1702e35466384b9d

Observation 318485fe-dbb2-4117-bac5-cd3e05544081 · inbound

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning cites this paper.

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning MiMo-VL Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:21.112448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:21.112448Z digest=sha256:e4898de407c9fd8b3712af749acb0a6a4e2f4c0985b2fb4a2ef108989c646062

Observation 1283b163-a25e-4989-99da-ea011bde4706 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MiMo-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.805838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.805838Z digest=sha256:9c50bbe5ea75d0f8ed8fce7d5a21cc2541abbf204aab9b0f3edf58931f59e825

Observation 83f477d4-a0ee-4e46-b838-1ef3770b3e50 · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report MiMo-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:32.039808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:32.039808Z digest=sha256:3ec2390726e05f4aecb0bb91708a3210eeb50add27f4f673423e7d3206f0e121

Observation 66398f4f-8c4e-4ba3-b31d-33a74d5d8454 · inbound

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing cites this paper.

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing MiMo-VL Technical Report

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:06:50.111502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T20:04:00.773443Z digest=sha256:bef97c5ce0cadbe613ef8ca78f7a97c0f2b3a9ef342e2a4f1bd65640754f6813

Observation 0c8d68e1-e4a3-49ab-ae44-cdf495f6c7a2 · inbound

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality cites this paper.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.883362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.883362Z digest=sha256:f486753ed32b2702047a3245c24ff4e5116f102ca07d73278d925cd22e2d8303

Observation f71a09d5-6d54-4b7c-8b97-ee11e5773ca8 · inbound

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation cites this paper.

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.595790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:42.595790Z digest=sha256:112a7e4400632d306cad8841991b79bb004acc1dccc641be853697051e960f93

Observation 297e34fa-3fdd-4af5-ade9-56901d741640 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models MiMo-VL Technical Report

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:57.154824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:57.154824Z digest=sha256:0389084a11b0a72f2c2e3adb3f410010441dc444244eee7a93eb2e278dcc063a

Observation b12fa6d3-9b00-4db2-95ce-68799fe4fb10 · inbound

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning cites this paper.

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning MiMo-VL Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T07:02:44.094748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:02:44.094748Z digest=sha256:8633c4e4e1984c664684901844043d213c9130a58a8650a453bbb6f73f2a98bc

Observation c6250233-8598-4bfb-9014-55b6ec3f5792 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report MiMo-VL Technical Report

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.687903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:6dc20d666a5113927e61bee2589c71f84dd01e3e2a11020575a0e2acb0086e76

Observation c82b5776-93cb-4903-8d1a-afe3ddaf8a8c · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation MiMo-VL Technical Report

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.378830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:38:49.075457Z digest=sha256:4223aeb1f9e27379cec0c26a735b09fcf68f80c0e06f228eaec53eb2ea815280

Observation 73ec2059-50bf-408b-a4d9-affbcfca6975 · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation MiMo-VL Technical Report

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:24.167937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:14:24.167937Z digest=sha256:06fc3602153f35c3f396b459a76d7efce7e5c284518fcd4d6f8705ac24d52147

Observation 4bd98a5d-fed8-4758-8c26-3f7e5d93553c · inbound

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? cites this paper.

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? MiMo-VL Technical Report

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:40:12.591587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:36:28.844714Z digest=sha256:25634b5ab2340880cbee606b3edb2504a0315f4ce521c286ff96838bc7d719d1

Observation 8628686b-b34e-440d-8232-120cf27510e9 · inbound

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? cites this paper.

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? MiMo-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:31:21.274478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:31:21.274478Z digest=sha256:6b92af44dcae28aed2e4e2fb3a9d724aae91567ae2ca9dc6cf26f347ab0cf925

Observation 9f537819-a78b-4e4e-8494-e6c407fc8e08 · inbound

Learning Self-Correction in Vision-Language Models via Rollout Augmentation cites this paper.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation MiMo-VL Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.930637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.930637Z digest=sha256:540be8cf5505df92c6a1a086a0b2d3b90b7d787a7357a7eef3811701a98e76cc

Observation 9a820de4-a910-4a8d-a76c-6440165bd6fd · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.621170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.621170Z digest=sha256:0ad9b694b2fe69545dc260d798c5dbc44a508e3008e9f216dcd6d7e46f958389

Observation 54be8578-9b88-423a-8e54-6368f2bb3019 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards MiMo-VL Technical Report

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:01.057349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:89d54482e3d16cb9e6ecb44fa6789d4eb38a566be52d9dc725626eb66c3ec995

Observation 2ad5634b-2940-4275-8fe7-d9223b08657b · inbound

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations cites this paper.

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations MiMo-VL Technical Report

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.855125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:23:45.814042Z digest=sha256:8f8a38edec325d1587bc4736823e02413404fd29ef33625e03e402a44b62c0dd

Observation 93ddcea5-acf8-4ffb-8237-018a3f76e2a3 · inbound

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models cites this paper.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models MiMo-VL Technical Report

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.615360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:2f06b8bea95f46459acc53862721812417b3ab860431a83864cfdaea138583e5

Observation ccbc11e3-b927-4d3a-8df1-943a753f759e · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model MiMo-VL Technical Report

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.261247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:3f005562b9c9826e1653af39f2159bcf8c9a469ef9c57b8679fdf00bb7a6eb43

Observation a0dd9a59-629e-4a98-aa0a-89287382ae77 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MiMo-VL Technical Report

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:43.902486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:f7148bfd6b326ca603b730b600eb0cd0db199fdcea0dfb320aaf84af59972d3a

Observation 77cabe27-51d5-40c4-aa38-b3a2361805c2 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MiMo-VL Technical Report

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.700497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:b0b0cc93eacfda062bbd314697a6ab46b8bd76c543decad78fa13dca890fe3f1

Observation c4354bf9-2a8f-4871-91f3-eb6334a7b9e1 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning MiMo-VL Technical Report

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:55:44.279701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:eb0cf8e7236271ca22fe76eb067e37c33928ab4876aad33e1f87eb559a7e1510

Observation 0f5b740a-da5a-4aff-b708-a435da384cba · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos MiMo-VL Technical Report

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:55.990591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:d0abaf0e90b439c5d6600971758b3981b5e246fd5a87754636c297ba89ab310c

Observation c894e30b-280e-44aa-b709-a1bfb5edb5a9 · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models MiMo-VL Technical Report

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:45:58.468198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:72026ed0c218d0fbc7220890b672eb6e8358cd1d44c189f1a97daa5a278bbc7f

Observation 893cbb9f-1030-4cab-b4cc-941550d89e86 · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:37:03.526884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:35:09.568623Z digest=sha256:7f9e706d30bbcab2be7ee01609d34fd8deac1d7591fee0440ea6d30ca486075e

Observation 3ac97830-2507-4a0b-98ad-ce6f04e24b6e · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:18:00.124492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:10:49.059686Z digest=sha256:b8dc540132a4c33046851f4dbbd0d676af2a60864b659b489c1b7be04da7fe0f

Observation 0569c09f-ce8d-4b67-80c0-7a74e654719c · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.597384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:58:55.817334Z digest=sha256:628dc5b8ba85347b495d58854f1a046708080a15dc3419b9797e7d1ca265d31b

Observation 031d66aa-8f0a-4a00-b416-864604493087 · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation MiMo-VL Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:46.445675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T22:41:13.164195Z digest=sha256:5f70e433f91e6259f311e7a9f714af4bb8401e07571c994e560f6edf47630665

Observation 0334a5b5-72d2-44bf-8c11-5bb24cc56331 · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding MiMo-VL Technical Report

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.412313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:91999956d82ab29f10a5e4786f6b609c59ec31d4e74541d649b6bea44f30f8fb

Observation 9e181396-0d43-4b55-809d-c86d586c3484 · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation MiMo-VL Technical Report

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.489809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:58:01.621488Z digest=sha256:cdb785b2b638792eecfd4b142822f07dd8fafff0daf507c02fb58457e9a7688a

Observation e9a0c321-523d-4464-b0a9-c3eeeeffb87a · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation MiMo-VL Technical Report

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.506207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:28:21.605646Z digest=sha256:30bdc779614e41456a19d7f780e4c405e563256d83d26c7873295341c178da72

Observation f91bb6c9-f935-4c34-ab58-92281d03fbf3 · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues MiMo-VL Technical Report

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:14:42.401670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:bbab70efefbe87ccf644ac649ec50bb739d721cbfd9f078536ef12bbccfad584

Observation 83e8f609-e1d0-4b50-9f14-5bc4fcd3faf0 · inbound

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding cites this paper.

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding MiMo-VL Technical Report

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:06:15.393088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T08:04:55.680668Z digest=sha256:43998935dfceecea88b2f2e7b840890ebbb75d18f09002a8987aedd3642f7e82

Observation 7fa44757-a829-4dcb-9436-49bcdd706b9d · inbound

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis cites this paper.

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis MiMo-VL Technical Report

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:54:43.935849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T13:51:47.455642Z digest=sha256:2432f81ba381bd5b0a9af655b6730c892db3eaede4ef924a2bf27ff3d8a4dffe

Observation a5a691c8-6b18-4d0a-91c3-a657e0f751e0 · inbound

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker cites this paper.

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker MiMo-VL Technical Report

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:02.166045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:15:02.455554Z digest=sha256:401320b1d538d360cb7f19f0e8d3046d05f1251399ac460666cda863000fd4c5

Observation a6ec13b3-b844-40b6-8ed8-9bcf9aa26393 · inbound

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning cites this paper.

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning MiMo-VL Technical Report

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:42:46.667492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:38:30.213948Z digest=sha256:1e33cbdbb5f21d8434962c7fe59f44170cc06704404fa103a0f1d17977baf74f

Observation a3b539c2-eae4-444f-ae51-e861be05b60f · inbound

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL cites this paper.

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL MiMo-VL Technical Report

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.971431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:55:07.045625Z digest=sha256:bb0a275e05ff77ce6a33b86aec2ad12f4c7db237994a282c133082776c8fb5ee

Observation af881701-d82d-4273-839e-3ef5afd8e92e · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding MiMo-VL Technical Report

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:27.944526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:7a5dc7a2740cc3d5f29c5b033c0cdb8baa2066fe92b324cef6beb4501f88a616

Observation 9ba9fe9f-ca95-473a-a74f-74f5ed8b7c7e · inbound

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues cites this paper.

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues MiMo-VL Technical Report

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:47.984436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T06:04:28.939248Z digest=sha256:28c465f7b243a9e50c09a26dee57a15cc5ac804a75e05dc50d24782b5ab6b7b0

Observation 229438a3-c704-4337-85dd-68bfdd291a7a · inbound

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding cites this paper.

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding MiMo-VL Technical Report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:24.811296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T19:40:56.851297Z digest=sha256:b7594deed989cc7288a47d546b79330c4712db20ad89b00a6ae965a3f726c74a

Observation f312811b-cd21-4be4-9a69-dd19e9c8f32e · inbound

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding cites this paper.

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding MiMo-VL Technical Report

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.033020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:04:29.739632Z digest=sha256:a6a574f3a8267a91b12bd4e38007f4002d6f0c8fa3c0e772768593d42c9c23aa

Observation e8f54eff-88c8-4653-95b8-f558789e19cf · inbound

AIR: Adaptive Interleaved Reasoning with Code in MLLMs cites this paper.

AIR: Adaptive Interleaved Reasoning with Code in MLLMs MiMo-VL Technical Report

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.983706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:06:38.001604Z digest=sha256:f1df6c46df1f553973e248775ef2fd6d1ff3110c0226da1fa87ac5b3681cf2fc

Observation 6a3f5543-80b6-4747-bf82-9c2e3d3ef69f · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning MiMo-VL Technical Report

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.283757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:775c6d36df59a72c197349db6c7fb8e8662bc44b66218f3984548ce64fc8b617

Observation 991bc27b-06fc-4768-8e66-36a67bbedc2e · inbound

Aloe-Vision: Robust Vision-Language Models for Healthcare cites this paper.

Aloe-Vision: Robust Vision-Language Models for Healthcare MiMo-VL Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:57.944069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T02:02:47.472868Z digest=sha256:5f60c34f4f5fa56430578f95643176c4ff8ba5f8fa7c6d369fdae218c3896d78

Observation 36501310-a408-4f0d-89a1-e547603e6e68 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models MiMo-VL Technical Report

Reference 277

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.594634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:30227435c4d5d62cbf4c8fc13678772807be8f35978d32984cf3259ecd972cc9

Observation 485ee3cd-1f75-48bb-ade0-16c332cb0008 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models MiMo-VL Technical Report

Reference 277

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.078871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:816eaa5514937e400208a2c4e5aad7a40034a21acfba2f5063931aae027f694f

Observation e3295a0c-7ae7-45dd-8c0d-f85951fb995d · inbound

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension cites this paper.

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension MiMo-VL Technical Report

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:48:35.007245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T15:41:03.548799Z digest=sha256:7367c96ef31cb79f09fcf2a994a633547ff7cfb542330925d2fc64e241b3bcfa

Observation 45c10caa-2eb0-4a62-b485-f74be4e92864 · inbound

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools cites this paper.

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools MiMo-VL Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T07:02:55.388823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:02:55.388823Z digest=sha256:83265a991417c692452ef5784ddf3b7e6dd36bbc5880975729631c4e1a7656f9

Observation 1fb177ae-653c-4280-b4d2-86794a0e404f · inbound

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence cites this paper.

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence MiMo-VL Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:30.687572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:30.687572Z digest=sha256:7ec19b5c850ef50be3db087907d5d96aff5af9b7dc5a44282e61b96b10ea4a35

Observation 8c9a65e3-14b0-4b3d-bf4e-54edb7612293 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs MiMo-VL Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.011960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.011960Z digest=sha256:f179781ad4f6a50df4349c548c50722fcb4c43c6252be505575246e1856444ab

Observation bc403541-2b97-4b33-9b74-c58a6aa4d875 · inbound

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement cites this paper.

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement MiMo-VL Technical Report

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:43.966698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:43.966698Z digest=sha256:43f6f976fd9d928fc763701521fb4b81152c00a9670d00995dcadf3a0680c239

Observation 38a6600d-e2b8-4c86-bd0e-470aeb0b9fee · inbound

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis cites this paper.

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis MiMo-VL Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-30T12:15:47.193300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:15:47.193300Z digest=sha256:959990401fa296f45aab416c22dd38faa18c7434ab3cd034935364ccbc7b4063

Observation 72382549-ea31-40eb-a9bf-70cb567513b8 · inbound

NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model cites this paper.

NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model MiMo-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T09:49:33.571423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:49:33.571423Z digest=sha256:60ca3bf66cf1122c50f6d55200e088d790937699665363ad0aa6a45564296706

Observation 958d1bf7-dbaf-43b6-8f45-47cd133a3c08 · inbound

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research cites this paper.

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research MiMo-VL Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:16.891205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:16.891205Z digest=sha256:9701149b83ddb7f32db709ac4a33a26fdf9a939a8705623d9e1408084787145e

Observation aaca9b6d-975b-4fba-a7c6-4d74fcfd061d · inbound

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes cites this paper.

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes MiMo-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T01:14:37.781329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:14:37.781329Z digest=sha256:7846d08c2bf590911447e394d8cfbd667804b6219b48e38612624cb52ca94ab2

Observation a7a12be9-d250-4a05-8de4-ccfca55f846b · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning MiMo-VL Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.095639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.095639Z digest=sha256:f63c59eb4996b594f753e8ac00895ba2bef21f2c5669202a9daff0105baec15b

Observation 1c0ff46f-77e5-48d1-af0e-05b385705b8f · inbound

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding cites this paper.

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding MiMo-VL Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T15:43:56.781142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:43:56.781142Z digest=sha256:836b923d21225fe777bf17a701ce3b68965b418672e9059c9a936b0982da89a0

Observation d5d883eb-ec4d-4405-ad37-cf7b0f03d325 · inbound

OPD-V: Visual On-Policy Self-Distillation with Modality Balance cites this paper.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance MiMo-VL Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:15.205510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:40:15.205510Z digest=sha256:df572627c5f637d1d7301b98757e25a0f49688bff534c33faf0a8831a6a5c4a1