Pith. sign in

Paper Citation Record · LEDGER

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

As of 4 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2604.03307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.03307 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T23:57:47.657243Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T06:45:27.857034Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T13:56:19.173208Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact26
  • verified fuzzy13
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56382967-f85d-4383-bfc7-4f1fdcae5cc3 · outbound

This paper cites Qwen Technical Report.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Qwen Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.490776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:f056aab2350930bd6ccadfb2698b92cadfaf3d7a2cca3e2fc17bb8508e7bc591

Observation bd5e5be2-8470-4c5c-822d-3b59970544e4 · outbound

This paper cites Qwen3-VL Technical Report.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.621212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:d33f32ddc7da5a6e8e42b800835b73f441769b404d8d28204347e09eb9e9c781

Observation 9da4abac-eacf-4abb-8323-84cf38282d8c · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.519794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:e6749cbcb662df02bd9481fc911d408f31010bd97ce7e52851356a7b93dc8d8a

Observation 2c1c0388-0b5c-4c06-b305-3f1797da3633 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic 11 visual-linguistic tasks.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Internvl: Scaling up vision foundation models and aligning for generic 11 visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.519461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:c0ee242b739b9d39ccf82cc71177054dc33a3a37848c68f0d039b754624d2119

Observation e77e0b81-a7ff-485c-a4ad-ea6dc231f864 · outbound

This paper cites Compressed Chain of Thought: Efficient Reasoning Through Dense Representations.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Compressed Chain of Thought: Efficient Reasoning Through Dense Representations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:47:40.470022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:ae30127281f3ed054b4d178db9537b0aad797ff4b10960040531476be9ce271a

Observation 605c65a9-2d20-43ba-beb3-2ee3961beb76 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Emerging Properties in Unified Multimodal Pretraining

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.578934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:50b178e4c2cf28075c584d72d4d57ff118beb3a719e488675b6eccb2a2b6311c

Observation 6a5d5814-3ba6-443a-8b76-c8c769a3f8f9 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Blink: Multimodal large language models can see but not perceive

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:d3d85a3f455149cb027cb766b49ea75051ed1bd8af12a676367bfcf7cabc3e41

Observation f54922e4-55e5-4b2a-ace6-55795e4ea7d0 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Training Large Language Models to Reason in a Continuous Latent Space

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.634771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:8c8f3bd68e37398c83027fc8c49839763f306ced42212d89104f9243e1fc1e7c

Observation 5aef65bd-63e6-4570-85c2-1007937c2f20 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.536112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:b0b00e1efc7246ee7ea64842e7c40549d9d4bebbf233abee9e8d972a866214b5

Observation 97d476b3-748c-4412-8e97-f641cc46561a · outbound

This paper cites GPT-4o System Card.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators GPT-4o System Card

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.430718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:a1da3067af096f959be8596f8f83d131463db62ebd73092b840f37baf46fccfb

Observation 1ef4b339-5566-48bd-bcdd-e477fa0ed9ea · outbound

This paper cites Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.440054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:addce9205b656b5f5fa046c82f9a7ee90f5c325dc570355e970ef8c16331fc51

Observation 1fe3b71d-7ac3-4323-867d-fa0b304081b6 · outbound

This paper cites Latent Visual Reasoning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Latent Visual Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.500607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:5834288dd9d17a63fd76e723e327a824d99fe72e07b03722d942e63961856eba

Observation 59493c6e-3948-4dd4-99fc-9ad9a9619ef9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators LLaVA-OneVision: Easy Visual Task Transfer

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.627283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:7b69efe7406cb2549579d1be8b7a44a47e4866ab3a0ba0a8f38ce855694dedc8

Observation c7782cb0-5b4e-40a7-b348-15a634e7e941 · outbound

This paper cites Visual-rft: Visual reinforcement fine-tuning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Visual-rft: Visual reinforcement fine-tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.576437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:0495bb9467e5ab0ff8d5768ec7a36e925979d59f9270b1755b168d69abffa116

Observation 8e5f7f26-dd15-4346-9226-aca42d9c3d33 · outbound

This paper cites Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.614087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:7d9d5e0b0be6eddba1c82768dc81e295f39a0b7b088623b1f9e7e30d8e930589

Observation 4ae5fe2d-3889-4424-8438-20a3013dad52 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:15:46.589538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:cec41540c84563dae0fdfb1258313c827f3093b8a75ad783622a711b729492d8

Observation fe43b6ab-f6b0-41d2-9969-3d02a9f9dab1 · outbound

This paper cites Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.514256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:999e790d8d5959e3fc19a70fa52e3691a39a39926fe16c9c0fbef8821baf4292

Observation 498e3f9d-8680-4b54-a970-898071c42b21 · outbound

This paper cites an unresolved cited work.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-13T23:58:29.556548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:1bfb7d24877971c5fe834a02643cc1e103cff4ef0f54a64af45f1e64d69f5998

Observation fecd0122-ba95-4896-9037-0a37a5c66983 · outbound

This paper cites Codi: Compressing chain-of-thought into continuous space via self-distillation.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Codi: Compressing chain-of-thought into continuous space via self-distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.561411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:cab5269f356e2a578a3dd04c1f1e04f9e520bb7ff00c078a694186926df746b4

Observation cd2c0691-27e4-4e05-9179-be2baea0d56a · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:13:13.788897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:f3dea8b55a71718a552ef3d34a481ce16c54350fd3835132115b65cbc0433e6a

Observation 005aa765-86f9-47e7-9a9d-25355ef293c7 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv e-prints, pages arXiv–2503.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv e-prints, pages arXiv–2503

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.567637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:4085979ed07dbeefde8516874559d32605b481b3c64f59e7c533db7264c27a1f

Observation 4d3a1248-6fc4-4d6b-b1a0-49d4f7afe4dd · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.515196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:d803d15b7041c8d6971327168152d688c1596f1fb9db9b9fcd89e368832e6f80

Observation b4775fcf-8cfb-4ec1-b317-b1c75584d709 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:22:27.090901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:007c21e4c50bd935320763610535ad434b379a342a3091098fbf0665981fa329

Observation c06d79e1-8957-4c62-a998-f2bd30d29bfd · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Monet: Reasoning in latent visual space beyond images and language

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.543353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:09522dbcdcdce8bba8caf59d7ba1b183e36c1c6659c11957756b7eb12200909b

Observation 6285fb97-d596-4582-9ab9-a50195a27335 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.600076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:850e3f6b3caa889b1f3592fa92c7942069a07e3c6c62876696a19d636e3c511e

Observation 1a511570-42be-4fb4-9955-2269a7bc8953 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.544943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:26c08b617024f3fda301d71445cd9365c067aa97db31d4af5e395335682223fc

Observation 64d0fd5c-6c9a-43a2-8266-583407ae2783 · outbound

This paper cites Perception-Aware Policy Optimization for Multimodal Reasoning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.528420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:29d6897f37b897ba235b7f943841fea74a6d0e1f968435756557527d3666f559

Observation b2c3769f-d601-4f3d-8b57-f678d12c6beb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.549587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:a3e699bb173646e1afe06a350bf11802f9500f6a6e009d6fd8953f1500331b14

Observation aa29c6f0-164e-43a3-af49-7f78e3ef4753 · outbound

This paper cites Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.593993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:16e50b3843e3d6ce42c0f324e00e2d14e128bd065674a3b7b5a987d4767dd414

Observation 4a31d6bd-3e61-4c0b-abaf-38c030a5807a · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators V?: Guided visual search as a core mechanism in multimodal llms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.533713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:d61894f7f8608db346866943ab364aa53ec933926bc194eac305f0f38bcd85dc

Observation 54428619-bef9-4b63-b9af-7efeee0f72b9 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Llava-cot: Let vision language models reason step-by-step

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.538739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:bf683e6d5d04f71e320d9ddb671aebc255fd176b51eed2bb716964a85ed2f613

Observation 2640e6fe-3f95-44d8-9628-0d999270d5c9 · outbound

This paper cites Mc-bench: A benchmark for multi-context visual grounding in the era of mllms.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Mc-bench: A benchmark for multi-context visual grounding in the era of mllms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.523946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:17bb8cd96cb536c89ab81db525280eeebfffdd6795a13afc00779ab57e3ad763

Observation ce1f991b-3bb5-4683-84af-b9431f06f8b1 · outbound

This paper cites R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.528009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:ae3f4d30e12dac89bd023e98bbfcf4719af9632b33e546e9f2643fd0d0bf39fa

Observation 570e609f-3dc3-41ae-9f50-edcad274807a · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.561653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:d5670c551c7649c89f37a35139f5ab17cf1d7c8b3bf206b088b215138c5ef218

Observation 83b88eaf-9ed7-4802-aeb5-e0b6c180f8db · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.586668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:b790e54d1ec8a4a0457071deeee05f63f7d1c2929c3a71c048cbb64c9e77b6a4

Observation 35f11d4c-8ca6-486e-b4ba-a6a9c0554951 · outbound

This paper cites Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl.arXiv e-prints, pages arXiv–2505.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl.arXiv e-prints, pages arXiv–2505

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.572747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:cb2d05a674c8aeadef5b9bc8977ce824a139be49123c6109ca1a55ff22966114

Observation 2f98649b-1879-493f-8174-555efd4cfe0a · outbound

This paper cites Thyme: Think Beyond Images.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Thyme: Think Beyond Images

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:33:29.495992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:31d08dc5d9fd811f3e98b9e7936b8ca1ccc526a7ca6fa3a7922b9c454c8b7137

Observation aeba2d28-612a-4e41-8a8d-ec6c810ed36b · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.958879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:48f0f77db8488ddc0f94eab677a0463ce1449e3d09c9d0d58c1241b814d8801b

Observation 4695c34f-1a6e-4728-be36-9b5ebaee0ced · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.607162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:1fc8007f1638f19c127bfaad2b1e0e332e2a708f46734af526fa4df2e3e0e483

Observation e0907d01-1ba7-4334-a8b7-d8447b7d2acf · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.497797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:041ceed520353fca4f96dfe8da3f89388b59b3e3485aabf887ce33431edf4f3a

Pith citing papers

Observation d9cbee03-fca8-467c-b2e4-33c98b9a092a · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

Reference 86

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T20:22:37.724913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:36dd1c69e8fe1153cc5a62ddfebb6a8c76633f878a5e372e6f95f9797395698d

Observation c2ddf2ac-a0d9-4960-9452-93d250756ea7 · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T13:56:19.174675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-09T13:51:49.149342Z digest=sha256:2808a5fca4c6ca7060df9823462b25a4928fa93aaf4e55ca608becf8e1ffe7a2

Observation 0276d545-cc89-4624-9e14-bdd0a3a33304 · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-13T06:45:27.857034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:45:27.857034Z digest=sha256:a43fc1dbdc13ff00c0a5c76a45163ed750c73066b3357f6224cfe1b6c5fcde0e