Pith. sign in

Paper Citation Record · LEDGER

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 7 inbound Pith citation observations for arXiv:2506.22434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22434 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:23.293868Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T02:47:38.676727Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.937856Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8b68a97-0e72-4a9c-9ad8-c454e6497c4d · outbound

This paper cites GPT-4 Technical Report.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.661051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.661051Z digest=sha256:b721be1803db504c46ce5e41bb0205bcb8bb460a32363f2d67a7c1b70034a708

Observation b4c6f8d8-18de-4fd4-b427-3fca9ac4deae · outbound

This paper cites Qwen2.5-VL Technical Report.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.714681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.714681Z digest=sha256:7d4505bac274a8eb28e07f183cd7c3b109afa67a9cdbc9fea2c4b8721e3b329e

Observation 6e0fbd76-dce3-4de0-8570-97c3b8910b53 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Emerging properties in self-supervised vision transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:28.722552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:19.841634Z digest=sha256:81c9cc5f60d613385021d21ee60e3e9b8bea29fd3092d9fa86084afd605029a4

Observation 24113702-aa5e-43d8-9559-237ebc7b9ba6 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.944658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.944658Z digest=sha256:ed96d75406cfd56358b58e15f19f2414cab11aed4a287915b44c6e24a8a3deb5

Observation 295ed156-2397-45a4-8b99-322dc2aef68e · outbound

This paper cites an unresolved cited work.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:28.465181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.024347Z digest=sha256:60b3a210f60539c212b893fd1106a35da2140fee9e22a90ef8cf9c17c3dc68f3

Observation febea6a5-eb8a-4871-b738-a26a0ff716d5 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?NeurIPS,.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Are we on the right way for evaluating large vision-language models?NeurIPS,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:28.130827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.104973Z digest=sha256:a74a31bf95d0f0a543498667af589e8ef1041d6d852381b462b726037a2521b4

Observation 1ee3d4c4-9c61-42d0-b067-2fad28420e2d · outbound

This paper cites A simple framework for contrastive learning of visual representations.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning A simple framework for contrastive learning of visual representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:27.798641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.165768Z digest=sha256:fca911944062d541556a923df1027ed3e6b739d05407060ce156be892712d2bd

Observation efe40d4f-71f4-48da-bbb3-9fe78dd576c5 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:27.481876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.214041Z digest=sha256:c6092cab65711de14ba636fbc136602da14354ae6ffa4e38c065e1131a42e44e

Observation 35e6b1cf-aba6-497f-a8ab-26cd02e18cad · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:27.260370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.240390Z digest=sha256:0dd37fe6eb21df50e2c1de584331817aabead3150d38d1d8cd560bc5e7937b74

Observation f806f609-79e3-4de7-9241-2314288d1859 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Blink: Multimodal large language models can see but not perceive

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.991239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.266790Z digest=sha256:d60983958c1ef6f157fda3bb190b539f5bb4a618d5c6b6071137b29efbcc9741

Observation df7af8c2-2dd1-4a87-9d21-0f599b0a5c6a · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.661498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.284666Z digest=sha256:f395431b73bb2dc083e7a2ec6774904a02491197bf285736809024bbbbcd3ad0

Observation 6ab5d93a-f684-4a9e-b584-7c5e5b073eec · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.303571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.303571Z digest=sha256:cb7e952095aacbabef8fb3cdb6b22963c2107645569d45a3d74d299d62f8fb59

Observation 3ad90e29-c3eb-4a72-b5e7-8228ec987354 · outbound

This paper cites Masked autoencoders are scalable vision learners.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Masked autoencoders are scalable vision learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.308149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.347365Z digest=sha256:1e0ae8aebf67e042deae7638337c5db39413e75a2e70662410b81e7b03816e17

Observation a1702eba-3960-450d-8ef9-9564231be04d · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Momentum contrast for unsupervised visual representation learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.002413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.380458Z digest=sha256:2ffd93abdc384a8c4962b6e503fefbc28bc76f2ff64b151d3209d59bfce2310d

Observation 30c8b7a1-2087-4abf-9a58-30f83bf48927 · outbound

This paper cites GPT-4o System Card.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.418100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.418100Z digest=sha256:35873607422b970a945b153ef2b40e2852443975fca2c32445a5882c16da0a9a

Observation 3e64c7b2-2243-4638-a3ce-b61883e05687 · outbound

This paper cites OpenAI o1 System Card.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.487532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.487532Z digest=sha256:e48e0e16682ff7778f03b34d38ae18a851ee192c397c7f9a1921280ec70fbd66

Observation cff9c650-f687-4f86-9f3d-cdb099491a55 · outbound

This paper cites Llava-onevision: Easy visual task transfer.TMLR, 2025.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Llava-onevision: Easy visual task transfer.TMLR, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.699586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.636213Z digest=sha256:4378cdcf4553b52a26b64b7262f6a6ecc0ea16c29eb5b85d9f71917a0ea9b262

Observation 491e6990-29ec-4934-a82d-bdedd27f0979 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.772845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.772845Z digest=sha256:66693fd789fbef6db09d37551fc91c093a7d555f3ed67925daa784a206a84f9e

Observation 694b5750-16b2-49ab-b776-38b59ea66a5d · outbound

This paper cites Visual instruction tuning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Visual instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.567251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.889944Z digest=sha256:d40beafc370642ba039a67ca07c3a053eb8521d31d4c76e052aeddd04a59b06a

Observation dd08eda1-bfeb-49df-bcf4-1b3c9cfd4720 · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.arXiv:2504.13055, 2025.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Noisyrollout: Reinforcing visual reasoning with data augmentation.arXiv:2504.13055, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.936461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.936461Z digest=sha256:2434fef42770d3ee1680958743dcfe940241e79f5e74f44559db97de655b198e

Observation 0d8b729d-bc4d-4c58-800b-669a16e0a869 · outbound

This paper cites Mmdu: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for lvlms.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Mmdu: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for lvlms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.436256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:20.990374Z digest=sha256:af608a75f72f3c46ab5d8efece1af180900e802b40dc762f6d483f21646e56bf

Observation 6def041d-365c-42d8-a439-367f88f6ffc7 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.298666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:21.089298Z digest=sha256:9328243238c3ddef10655d2c45e380b627b66721ed1c58a1e7e6c220d5180c84

Observation c2810585-c011-4c7d-8ccb-bc2427e70979 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.182215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.182215Z digest=sha256:29bcc54b36e6f7db86950460e9c754a618897c39daaa19b7d6cfccf51d0b7212

Observation 508e4ff3-66f7-41f2-a4d8-e6703505eadd · outbound

This paper cites Mmiu: Multimodal multi-image understanding for evaluating large vision- language models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Mmiu: Multimodal multi-image understanding for evaluating large vision- language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.162989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:21.303114Z digest=sha256:8d95d364c458cfb12eff7a829050da0fcdf28081bece1fdf2cb8697d7c42ecb1

Observation 41ca8a4b-5b82-4ff3-a3b4-f32a305a2bd7 · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.421604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.421604Z digest=sha256:d3a09d6e72715b5b5dac4701a0af30885f116d2a8e132208de849ea10604fbe0

Observation 1f566777-8f5c-4265-890a-d60523701d0c · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.546619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.546619Z digest=sha256:f65b2367b774005c28aa40eb3428bc7f5a06aaeb43438fecc2cc401ae4ca91f6

Observation 74b28711-97ec-463d-9e80-aa6dcca57847 · outbound

This paper cites Seed-thinking-v1.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Seed-thinking-v1

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.645213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.645213Z digest=sha256:16d0d96ec0b5d741f988dbc1e24f72995e5f0a6cc4f42bd7c21debe5ab41ca1d

Observation 429e076a-6e03-4d5e-b0b4-22dd59d2653b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.712099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.712099Z digest=sha256:03db087053db724cae644b7e103feae06136936fe5d50f7f1b4c7e19c85fffa4

Observation 82cc9a97-fd6c-4e9d-84df-525ab92e2c2c · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv:2503.20752, 2025.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv:2503.20752, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.814227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.814227Z digest=sha256:afd50ef6945507104073a85f7927fc9e2313556f58f682bc25ba742ad3513946

Observation f858776c-a47b-478c-b409-976b86053609 · outbound

This paper cites VidGen-1M: A Large-Scale Dataset for Text-to-video Generation.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.932155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.932155Z digest=sha256:b62d08a1db760eeefe38de0886a9a447a202af140a8ecb2928438956b6656d33

Observation b5ccabef-a66a-42c6-a438-ff63745e2410 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.028060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.028060Z digest=sha256:e7a0895359a880b969a7cb4d7792f1e2b18c5b64024c6284fffa7984dbeb9f4c

Observation 7d9b5079-b267-42ad-8466-aa1573e257a6 · outbound

This paper cites Muirbench: A comprehensive benchmark for robust multi-image understanding.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Muirbench: A comprehensive benchmark for robust multi-image understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.033961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:22.089405Z digest=sha256:4d71d30e5bc87cb680a29af3652524a2b24dd21910aed3ba3a8967b78685adaa

Observation 65027c1c-7b21-48d4-9e46-8f1b6da2f3e1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.153445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.153445Z digest=sha256:9f076fcd114f5144be7fa7dc0ea921a904ac36862218cde5ea6afd2332104844

Observation fa107860-b77c-457c-bd26-239bb6829ccf · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.226382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.226382Z digest=sha256:b122676f8f839efb75f1a088a5a707e828edef46da914eece83c4cac10a75d8c

Observation 89425d77-f32b-436d-8a79-d4d75a0b76e4 · outbound

This paper cites Omniedit: Building image editing generalist models through specialist supervision.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Omniedit: Building image editing generalist models through specialist supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.899197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:22.310213Z digest=sha256:a53ed3e290d2d688f08dd69629831552d12605a175d883f3f9fe0abdf88dd066

Observation aa15e761-d7e7-4dfb-bd56-d4c133797abc · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.792525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:22.379211Z digest=sha256:e0466f0356099ef1f015fc0bc6270efa6be5dd064c2791e99e79bfd9ba3694a7

Observation f311559c-40f0-497b-bd9e-f8c5f4fb16d1 · outbound

This paper cites Towards open-ended visual quality comparison.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Towards open-ended visual quality comparison

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.690395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:22.476519Z digest=sha256:618d0f0e75dfb53b2af43f845168be5bce2a877b5ce3f9c43b4cdc63818bfc95

Observation 7a41f29b-b701-4235-b643-88ca56801731 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.556572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:22.548940Z digest=sha256:72c10e1bf0067fb9e290bd86620032f74cf5460ae43a2cba7416c6053c2aa2ce

Observation 90f9bd71-b043-4599-ace4-89f419051ef8 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.453840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:22.635408Z digest=sha256:4d9d37e34860046f7eaced45a24b2b259bd386b3c0f6cfdaf28613a24c660810

Observation 0e367bb4-2cbc-4d97-ac8b-e3719455b251 · outbound

This paper cites VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.722405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.722405Z digest=sha256:534c6c9ee8d7695af8ca437b6283ad8166b6682f2e1b0d5559aed235a916b429

Observation 1ed0821a-b305-4cb4-a1fe-aa014113d41e · outbound

This paper cites Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:10:23.589277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:22.804765Z digest=sha256:1fcd45a2888df108cd76c0c462d33cd28bc82d8b14dfef693c9ea3444679a254

Observation ac77033f-768b-4197-b7c0-04d0923444c2 · outbound

This paper cites Long Context Transfer from Language to Vision.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.903154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.903154Z digest=sha256:b42ce93f298734e068d4657df6dbbc0593c997062f0657526efd180ad9a808e8

Observation d6e43932-032a-497b-a74f-48c42b77f00f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.971564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.971564Z digest=sha256:053582c038f22441991269369980200a446256083071f38f34b13ebd8dca2101

Observation 8ba2266b-47ea-4432-87f9-5ebfd913950d · outbound

This paper cites Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:23.068331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:23.068331Z digest=sha256:d4a0c034df65e2b1edcd178f60d568f36c2bcb7cf83d3635f3c13c53f7f8fa7c

Observation 626392d1-1c8c-4849-8999-a0fbb297730b · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Ultraedit: Instruction-based fine-grained image editing at scale

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.301168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:10:23.158141Z digest=sha256:f6ba7c293086fccaa051f22189cf8bd39a7689a4c4da8b9b849abc05470ca8f8

Observation cab53f23-af9c-4c36-83a5-25d31c6b7cc9 · outbound

This paper cites Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:23.293868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:23.293868Z digest=sha256:aa94bfe43df41895636ae81cd651d5802023e871b644df614eda1a2ab21a830e

Pith citing papers

Observation 56bbfb6a-5856-4629-9ef6-5dd15c4e0698 · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:15.166907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:15.166907Z digest=sha256:8f50de106e29cfbc2455cd984a06ebe4305c7cd40dbf65c8418d65da7d9132ec

Observation eac79354-ed08-4554-9934-83e1f5b37637 · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.367742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:ea6ae534222beb2e2513ac3f72c55edbfd8bd0d24f32443d8c31b1af151a2c91

Observation 4ec5dc58-95fe-426b-8a51-454a61ac8dc6 · inbound

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams cites this paper.

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:37.943816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T14:24:19.702588Z digest=sha256:98572a76a989433ac88632ef3456a5d2b1f49012112e8e880b426698cba5b562

Observation ef3e5ca9-6892-4534-8b80-e6d8b0a38be8 · inbound

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams cites this paper.

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:38.676727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:38.676727Z digest=sha256:cb52d448179077c227cfe30daa8abe67656f34a2960e56dcada7d20d6c7bd214

Observation d0a7c20a-8dbe-416a-9703-83022d97f974 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.939662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:91fe4591ad0f18910b061727ac5b49704d1675cec0f5db7d2cfc6b0a896c92b6

Observation 5d3575bf-4f7f-445a-a8df-ca754d0c73fa · inbound

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning cites this paper.

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:05:51.451562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T04:09:32.397341Z digest=sha256:7987c7b49b208ed46c8ac022a7a68e141847d40bf390800d7a79dce98c8398b2

Observation 7c81efc5-ace3-4125-b140-de5c81811780 · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.360790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.360790Z digest=sha256:8ee967e5737f46cd0c4b631a4c1fa2a4a38ff595148fb8c736de45db314ca64a