Pith. sign in

Paper Citation Record · LEDGER

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

As of 19 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 7 inbound Pith citation observations for arXiv:2506.22434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22434 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:23.293868Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T02:47:38.676727Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.937856Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8b68a97-0e72-4a9c-9ad8-c454e6497c4d · outbound

This paper cites GPT-4 Technical Report.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.661051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.661051Z digest=sha256:5c1e7cfce1c26b143a5141c1d67539f61f5d69e2a523c729109669ddcc1e6b46

Observation b4c6f8d8-18de-4fd4-b427-3fca9ac4deae · outbound

This paper cites Qwen2.5-VL Technical Report.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.714681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.714681Z digest=sha256:49356b1a9bf47712e8c05cf249cad26ac36e043f72c2daf0bbc7db25dafee6fd

Observation 6e0fbd76-dce3-4de0-8570-97c3b8910b53 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Emerging properties in self-supervised vision transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:28.722552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:19.841634Z digest=sha256:03be815bd75e6ad11ddd5f6f2dfcadf9b3b6a7c0da9b196c2288f1e3fa9107ff

Observation 24113702-aa5e-43d8-9559-237ebc7b9ba6 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.944658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.944658Z digest=sha256:2ae0d4acc400fa3368d72e703f4acb9bf828f0bea384c9075a5fe5da224a4149

Observation 295ed156-2397-45a4-8b99-322dc2aef68e · outbound

This paper cites an unresolved cited work.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:28.465181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.024347Z digest=sha256:c8c65d35fe69b916ba01e274f81384f760d12c56ee30ed3e297da9a066fe6c6d

Observation febea6a5-eb8a-4871-b738-a26a0ff716d5 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?NeurIPS,.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Are we on the right way for evaluating large vision-language models?NeurIPS,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:28.130827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.104973Z digest=sha256:dac8711f6c0c47951d9c7368a517bed943276339fc7ba05dd191133b11bfc222

Observation 1ee3d4c4-9c61-42d0-b067-2fad28420e2d · outbound

This paper cites A simple framework for contrastive learning of visual representations.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning A simple framework for contrastive learning of visual representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:27.798641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.165768Z digest=sha256:a6fb9478867058e1ae14b55afe6cc49d8ce25285bbed66deec61364a509289ec

Observation efe40d4f-71f4-48da-bbb3-9fe78dd576c5 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:27.481876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.214041Z digest=sha256:c8eb51523e5ff20a9263751ec22e5b483ca0d0b206069446d7b0f6131ea564f8

Observation 35e6b1cf-aba6-497f-a8ab-26cd02e18cad · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:27.260370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.240390Z digest=sha256:cd10757fc2842267a6ec4f703ff614b43de6f488cbec52fe03f9ff5c4743f154

Observation f806f609-79e3-4de7-9241-2314288d1859 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Blink: Multimodal large language models can see but not perceive

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.991239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.266790Z digest=sha256:d821e46124e4718de41452562098ef0c78eb57d6768b9e0cda64b28a663031f5

Observation df7af8c2-2dd1-4a87-9d21-0f599b0a5c6a · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.661498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.284666Z digest=sha256:4b6ef086220dc9e11b825d8e3887f377af979e1dd898d9a4a0afe2b3184818f7

Observation 6ab5d93a-f684-4a9e-b584-7c5e5b073eec · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.303571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.303571Z digest=sha256:b747d49d396856fcb0d0150a233e2fc9e2b2296b42d2dc45701b95e180b5fba3

Observation 3ad90e29-c3eb-4a72-b5e7-8228ec987354 · outbound

This paper cites Masked autoencoders are scalable vision learners.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Masked autoencoders are scalable vision learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.308149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.347365Z digest=sha256:76f8f5edc6659b9d0305f5552e70b911f26bb63c0f26a484cf5bd26fe5bbd467

Observation a1702eba-3960-450d-8ef9-9564231be04d · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Momentum contrast for unsupervised visual representation learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.002413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.380458Z digest=sha256:30679a022e3967325aa67a107c9e8e35bd69cd72e5ce4afb82cc09c419da0c43

Observation 30c8b7a1-2087-4abf-9a58-30f83bf48927 · outbound

This paper cites GPT-4o System Card.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.418100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.418100Z digest=sha256:c8c84a7a2d3f8b1af0c78254fc6d77f167fa51c583ca4617656685495117959c

Observation 3e64c7b2-2243-4638-a3ce-b61883e05687 · outbound

This paper cites OpenAI o1 System Card.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.487532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.487532Z digest=sha256:efdc7f40a57e43772621643f10813331520c2c4880204eb21ed753594bb2104c

Observation cff9c650-f687-4f86-9f3d-cdb099491a55 · outbound

This paper cites Llava-onevision: Easy visual task transfer.TMLR, 2025.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Llava-onevision: Easy visual task transfer.TMLR, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.699586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.636213Z digest=sha256:969668e28c1d852d4584fd0f8311eb97f26885789bb6875a426e31eb6ef7b950

Observation 491e6990-29ec-4934-a82d-bdedd27f0979 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.772845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.772845Z digest=sha256:fff745463b79f44da856b562e1b87915fef460f5be30d5ff9386a7ea031e7683

Observation 694b5750-16b2-49ab-b776-38b59ea66a5d · outbound

This paper cites Visual instruction tuning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Visual instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.567251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.889944Z digest=sha256:e547375e27f2091b5c579dd4485127c08e49f76cbd58c906541215e4abd21f13

Observation dd08eda1-bfeb-49df-bcf4-1b3c9cfd4720 · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.arXiv:2504.13055, 2025.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Noisyrollout: Reinforcing visual reasoning with data augmentation.arXiv:2504.13055, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.936461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.936461Z digest=sha256:cc60c97016ae33802c1f8cb9e24fe78adf59d0641b630f0f172ac2c8defb8137

Observation 0d8b729d-bc4d-4c58-800b-669a16e0a869 · outbound

This paper cites Mmdu: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for lvlms.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Mmdu: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for lvlms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.436256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:20.990374Z digest=sha256:e45a5789bba31dfe828837da7b92557e316588003c4f61a1a5faff8864ae7b76

Observation 6def041d-365c-42d8-a439-367f88f6ffc7 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.298666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:21.089298Z digest=sha256:c441a76b2a5cce268d8f2caff5bf1df80aeb8cc8f5a0d8979c8a2fe79ad5f00a

Observation c2810585-c011-4c7d-8ccb-bc2427e70979 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.182215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.182215Z digest=sha256:b6b6c8f86ef23133260eb6a50abb36125d6dbd73be61de8800b60b20a879d99f

Observation 508e4ff3-66f7-41f2-a4d8-e6703505eadd · outbound

This paper cites Mmiu: Multimodal multi-image understanding for evaluating large vision- language models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Mmiu: Multimodal multi-image understanding for evaluating large vision- language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.162989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:21.303114Z digest=sha256:3db9ff6e947519a56c64992a03290602803a0d6cd1d7d3328aa69aa9ffa94d13

Observation 41ca8a4b-5b82-4ff3-a3b4-f32a305a2bd7 · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.421604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.421604Z digest=sha256:eb7148abf51282a85f762e4d54a738df12f8d3684a8e1af285c4be42fa3e9825

Observation 1f566777-8f5c-4265-890a-d60523701d0c · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.546619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.546619Z digest=sha256:8c12f7e5c9117a8300369fa543a127606be99e1a865d9bb243420da3d108d9a4

Observation 74b28711-97ec-463d-9e80-aa6dcca57847 · outbound

This paper cites Seed-thinking-v1.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Seed-thinking-v1

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.645213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.645213Z digest=sha256:6cd92ed0acf4d6b296341d27c59f805400334e30d144bd41476f7e499f6a9b49

Observation 429e076a-6e03-4d5e-b0b4-22dd59d2653b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.712099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.712099Z digest=sha256:35a7416a1fb6c98e45535121cbdcf1d0c9a905837bd99862bfc1fcf06e60b526

Observation 82cc9a97-fd6c-4e9d-84df-525ab92e2c2c · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv:2503.20752, 2025.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv:2503.20752, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.814227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.814227Z digest=sha256:dd0c618c9c4e2fc3d7fba08aa0a383f4e8b89d8a4ab192f5a48b7b7116734e56

Observation f858776c-a47b-478c-b409-976b86053609 · outbound

This paper cites VidGen-1M: A Large-Scale Dataset for Text-to-video Generation.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.932155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.932155Z digest=sha256:046bf9497905db0fd409cc597e843bb48c90a89706d346c399e8b13688b2bdff

Observation b5ccabef-a66a-42c6-a438-ff63745e2410 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.028060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.028060Z digest=sha256:a8bcebef61335aab240dcedb7c7342d6cbfe0105809715de38fab0933ead6b5d

Observation 7d9b5079-b267-42ad-8466-aa1573e257a6 · outbound

This paper cites Muirbench: A comprehensive benchmark for robust multi-image understanding.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Muirbench: A comprehensive benchmark for robust multi-image understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.033961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:22.089405Z digest=sha256:2980d76f35be87dc8d0e5021376d60c94723a43b3ad4f68d1be5a05ec9092677

Observation 65027c1c-7b21-48d4-9e46-8f1b6da2f3e1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.153445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.153445Z digest=sha256:87302cbb2ee7e8e2ce86b050bc5ec69fff1871d894a03d49aa7449f953923a15

Observation fa107860-b77c-457c-bd26-239bb6829ccf · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.226382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.226382Z digest=sha256:c8ccb433b48a26e399c14178002dc88d0db70424a11b36a032a19307d9afc7c0

Observation 89425d77-f32b-436d-8a79-d4d75a0b76e4 · outbound

This paper cites Omniedit: Building image editing generalist models through specialist supervision.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Omniedit: Building image editing generalist models through specialist supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.899197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:22.310213Z digest=sha256:6bd1a2e3f48452a605ac329a33de968fb388997842be6cd2c57be47621d3f0f1

Observation aa15e761-d7e7-4dfb-bd56-d4c133797abc · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.792525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:22.379211Z digest=sha256:35a314821a212b07c3b6d9b2e8bb325f0afb186b45110aa16050cf54578be22f

Observation f311559c-40f0-497b-bd9e-f8c5f4fb16d1 · outbound

This paper cites Towards open-ended visual quality comparison.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Towards open-ended visual quality comparison

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.690395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:22.476519Z digest=sha256:dd6ff14db40f4829cc083b4e0a74255b89087d45dd5c93a902087dd267e22be0

Observation 7a41f29b-b701-4235-b643-88ca56801731 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.556572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:22.548940Z digest=sha256:4a03eb24d4bdae86bc83d9bea3ef6ac019a351f5e418de64a3b9d2c98ac27882

Observation 90f9bd71-b043-4599-ace4-89f419051ef8 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.453840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:22.635408Z digest=sha256:2a3ae1c1b7f774c9987781983b89e3b614105834e47f5c9b5fee25e78e788004

Observation 0e367bb4-2cbc-4d97-ac8b-e3719455b251 · outbound

This paper cites VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.722405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.722405Z digest=sha256:d524f56fb2daa4d8865945e12a86e605d90b2c4222933b67b62532e182f846ca

Observation 1ed0821a-b305-4cb4-a1fe-aa014113d41e · outbound

This paper cites Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:10:23.589277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:22.804765Z digest=sha256:87dc4ed79b73bbaf53c455fc3027950c08b4f42eb468ed51901fe58fda10c90c

Observation ac77033f-768b-4197-b7c0-04d0923444c2 · outbound

This paper cites Long Context Transfer from Language to Vision.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.903154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.903154Z digest=sha256:24760536250e2075d86e734c1552a4f33b26ea6bcbd7dcce726e0e83cba4de19

Observation d6e43932-032a-497b-a74f-48c42b77f00f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.971564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.971564Z digest=sha256:783783c11c2b591bce64dd2cacd9a9221c80d083f6cc2d114431546523c78d59

Observation 8ba2266b-47ea-4432-87f9-5ebfd913950d · outbound

This paper cites Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:23.068331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:23.068331Z digest=sha256:4fb6685f06b2ae76ea7a7c0c4825580ec883defe481a14de7ceb0f5c3e47a2c7

Observation 626392d1-1c8c-4849-8999-a0fbb297730b · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Ultraedit: Instruction-based fine-grained image editing at scale

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.301168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:10:23.158141Z digest=sha256:04c711b6b70c285d1ed8980a168ce89756ed06742b38ffb56ae48eacdbf2b018

Observation cab53f23-af9c-4c36-83a5-25d31c6b7cc9 · outbound

This paper cites Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:23.293868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:23.293868Z digest=sha256:3b756c6a568d3dec26e7d54bb590f751b178389e185d557e8a04df3674756835

Pith citing papers

Observation 56bbfb6a-5856-4629-9ef6-5dd15c4e0698 · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:15.166907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:15.166907Z digest=sha256:417916eb9b600b98c46c0c1e3ea731094b9f7da362063b540e1d98dbd0a1a608

Observation eac79354-ed08-4554-9934-83e1f5b37637 · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.367742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:83ce466c895610ccf80ea23151e82023230f593b58263f0c3d16fd3cee02fd7d

Observation 4ec5dc58-95fe-426b-8a51-454a61ac8dc6 · inbound

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams cites this paper.

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:37.943816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T14:24:19.702588Z digest=sha256:e68634c20299198867b2cba5694f819fb9f75b5a40384baf4ce02743b698b619

Observation ef3e5ca9-6892-4534-8b80-e6d8b0a38be8 · inbound

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams cites this paper.

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:38.676727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:38.676727Z digest=sha256:df2f19e1a130441ccecdb99109211cf7588064006ed163024ea76efb9635ea18

Observation d0a7c20a-8dbe-416a-9703-83022d97f974 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.939662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:9c578c4ce20c805253ffdc7f171ae747f40fca0b3e992f844267cd4a40e197d0

Observation 5d3575bf-4f7f-445a-a8df-ca754d0c73fa · inbound

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning cites this paper.

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:05:51.451562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T04:09:32.397341Z digest=sha256:822b62a388214e0acf755ba1c37bd39d531d185426c64ec9fe4892a6f21b8795

Observation 7c81efc5-ace3-4125-b140-de5c81811780 · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.360790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.360790Z digest=sha256:d06effb2d3befb6152f4a982dba59899bc752f21326abbcf38068eb08edd1297