Pith. sign in

Paper Citation Record · LEDGER

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

As of 8 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 8 inbound Pith citation observations for arXiv:2506.01663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01663 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:57.963146Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T13:24:17.538850Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.338873Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45c38f8e-d3ce-48bd-ba4b-8707c222c4be · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Phi-3 technical report: A highly capable language model locally on your phone

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:15.910483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:47.327341Z digest=sha256:7000242be54c66ad92a2ee2c79325fbbbd7fb293e906a5667cea8a18064590fb

Observation 21388ee8-55fa-40a8-a688-794109a643c0 · outbound

This paper cites Qwen technical report.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Qwen technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:15.658076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:47.436133Z digest=sha256:05909597e2ab71cf1473b2adbf5226079c4a2732f7d3f13f9db0166cec7b0787

Observation 76b437c0-1477-4cef-99da-25e08e82863f · outbound

This paper cites Qwen2.5-VL Technical Report.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:47.496339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:47.496339Z digest=sha256:b1047f62ee1567207359c91e7e944d1d9c7100314833416241ddc62f84e17031

Observation 3692d56d-5932-4216-8d21-bc4b56c42684 · outbound

This paper cites DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:47.677291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:47.677291Z digest=sha256:7d176ee92e7dfbb7705f0598c29c1c78ef6c9899e39002206e2916bb75c1c169

Observation 334684bc-b081-41e7-ac79-568dc72da4c9 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:49.596369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:49.596369Z digest=sha256:e614d38ebc17f9b55dbc12da93b6f15beae57f8102e9a4c181f825dfea8d1568

Observation 4887c136-36cd-4ac4-abbc-6bf9f3025d34 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.790506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:51.790506Z digest=sha256:2b2add405c89f4fdacee969cc38aa61ddb8d1633c74d17d1a8ee2486f46894a1

Observation 665676c7-b8e3-4351-8709-731161ae88a9 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.883143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:51.883143Z digest=sha256:7128290cfa9a76b8e3f08d833448ff8fa29ccbb7df3e6908e36235aa80f45972

Observation 32bca65c-6b13-4103-9b7f-e8af1db770fa · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:15.286495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:51.999902Z digest=sha256:6b2adc3fdf76cf98f2ef9fe9613556938b1c0697a5e6daf13a2e6b64c0955dbd

Observation 3d43ec70-be5b-486a-8192-5cd6012f69af · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.101099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.101099Z digest=sha256:fc62ebe5e4223219f3a340a722a7af5f7ef3984ef37fcc40e9b91cc424d825d2

Observation c4d04e08-e5a8-4602-bb1a-6b7f10711d09 · outbound

This paper cites Chain-of-verification reduces hallucination in large language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Chain-of-verification reduces hallucination in large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:14.972620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:52.179476Z digest=sha256:7d1e44953a84e77f43d5f4aac3f5a913089c03ce51bfbff9874d4219b6784c46

Observation f61066cc-98dd-4a66-8559-fed13358bda7 · outbound

This paper cites Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:14.746866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:52.260843Z digest=sha256:c5544b8d6e65fef1bcfd7e4e308ebea7a368a197ec85a3224df11a36d33b3fef

Observation 0013e321-7f3e-43e4-b1af-7a9c19f08130 · outbound

This paper cites an unresolved cited work.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:41:14.577533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:52.330821Z digest=sha256:8916af2caffa62aa5f51cb77b80a587820b42aa2a1f90256292496bcdfa8a8da

Observation 7294c1b8-4b96-4bbf-b770-25084a9a1065 · outbound

This paper cites Active vision: The psychology of looking and seeing.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Active vision: The psychology of looking and seeing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:14.394762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:52.400563Z digest=sha256:f0c0857114c8cf4b019e22fe03c8fd2e576aa06884b3fd03c71d6494b9fdda28

Observation 7586c170-250b-4a29-ba4a-6c91de48c5ee · outbound

This paper cites Small language model can self-correct.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Small language model can self-correct

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:14.225832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:52.465294Z digest=sha256:1bfbfa4e9e6b3b257547c43492cae0935c50087ced58a85541d7b4ec7e403def

Observation d8daadd2-8b5b-4aae-a3e5-1c645ef670c0 · outbound

This paper cites Eye movements in natural behavior.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Eye movements in natural behavior

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:14.018823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:52.555428Z digest=sha256:d6a9f93eba4a298bd217517f8d9a4d6dfe48dd9bda2db62bb909e272e85e5df2

Observation c30b6367-8779-4ae2-a245-70b18d5a1b9f · outbound

This paper cites Human gaze control during real-world scene perception.Trends in cognitive sciences, 7(11):498–504, 2003.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Human gaze control during real-world scene perception.Trends in cognitive sciences, 7(11):498–504, 2003

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:13.874662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:52.777140Z digest=sha256:68da827a628f9eb058142f67d5aa08573ecb7730023e8ab78211ab2ce55185e9

Observation 33795074-5261-4f2d-a544-74a39440397d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.921977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.921977Z digest=sha256:d37d7dc50a48130d54aa77f2998960332da76e97943653240b008fbca94e5a09

Observation 25624c85-eccc-45a0-a88a-5a25ead4f616 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:53.021225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:53.021225Z digest=sha256:1ba0a09e13b6f493cb9f11bdf692b5f061f7845ab9bd9b583fae6e554d30e13b

Observation d3b73bc6-8def-4a2e-a70e-bd96ec54c53e · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:53.103982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:53.103982Z digest=sha256:ba4f6f127d3145e5d77a0cf12be1e5372bb2e29a9badea92bf0a4d7cf96097b6

Observation 6bd05da0-5767-45c1-a096-ec6091e15086 · outbound

This paper cites Ma- tryoshka query transformer for large vision-language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Ma- tryoshka query transformer for large vision-language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:13.683219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:53.214626Z digest=sha256:75fc1358959748eb34ba368d9dbaebc7e04dd0a8dae266e1abc7c6b956cdb2a6

Observation 5d96b365-9739-4cf7-8d02-28482aa219f5 · outbound

This paper cites Aggregate-and-adapt natural language prompts for downstream generalization of clip.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Aggregate-and-adapt natural language prompts for downstream generalization of clip

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:13.573799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:39:53.375408Z digest=sha256:944042d7fc855c9866d6257ab6b46817d3617971c4ea5d1622c20c9030889799

Observation 9b0380ac-9b30-4c9e-8040-fa3c11b07d07 · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Large language models cannot self-correct reasoning yet

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:07.103123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:07.103123Z digest=sha256:c186b13f4814441094070b1141f5be41eca0bbca02d0e1c43df380a18733d4dd

Observation 21a94415-319b-4bb7-a817-e02e381b4212 · outbound

This paper cites GPT-4o System Card.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:08.965384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:08.965384Z digest=sha256:b3dabaaa4157687c30d2e47e798a1d8cbbb18b10d2b895de8c39e121871f5b4b

Observation 0392638f-e796-4657-9068-5aea7dfe2cac · outbound

This paper cites Maven: An effective multi-granularity hybrid visual encoding framework for multimodal large language model.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Maven: An effective multi-granularity hybrid visual encoding framework for multimodal large language model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:13.335877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:09.529263Z digest=sha256:99dfc8f431f84137e675fe80576629f2fc10d2c91e3219233c50757b698cc02b

Observation 512ac1f3-7669-4d76-a8f4-3eb2c4078c43 · outbound

This paper cites When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:13.149890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:09.725443Z digest=sha256:d693f75680156808c57c98cd3df28d3c4d3c9819444f6909c10c7b28abff4810

Observation 72cb6998-059e-4966-8f19-ba19b769d6a5 · outbound

This paper cites Language models can solve computer tasks.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Language models can solve computer tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:12.926936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:09.823903Z digest=sha256:2c82b15b121908390ef498f98728ddc219cfdbd1b7416253017056198f84c8f2

Observation 287cf32b-7167-4fbb-94c5-96fe62a9704e · outbound

This paper cites Training language models to self-correct via reinforcement learning.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Training language models to self-correct via reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:12.711503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:09.886914Z digest=sha256:435355bbeaed1fe72fe0b26ef0d104c4eb2ad3f89d0e1394ab3c87ecf5c2a0f8

Observation 8b5917de-feb5-493f-a3f4-011f9e3d746a · outbound

This paper cites What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874–87907, 2024.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874–87907, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:12.451868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:09.909382Z digest=sha256:dd1b8243ae17f4b7cc0ddba6f6ad9790294ff40363d1b3dd53769415eea43ce9

Observation 686006e4-f2dd-40fa-b86d-8ae837c87679 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Llava-next: Stronger llms supercharge multimodal capabilities in the wild

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:12.244196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:09.939325Z digest=sha256:45e5bc426f9c64a40cfbbcd55ce29d3640b70e14d06e8279a704db9a34976c7c

Observation 4634c676-34a9-4239-8107-53d9cc6d8430 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:10.032358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:10.032358Z digest=sha256:0c49d0ca381b69dfb95d5497496ea40cf1a0f441cc4b44053b884241aa86042c

Observation ca30a7ad-6ad2-4fd2-92f0-9978ef0cbcd1 · outbound

This paper cites DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.502823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.502823Z digest=sha256:f2e3100e0e31adac329557936d0d7f11ff8c08d4dc52533c643ad24fedae55c1

Observation 4a47c34b-7eec-4c16-82b3-8439b44816ff · outbound

This paper cites Cumo: Scaling multimodal llm with co-upcycled mixture-of-experts.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Cumo: Scaling multimodal llm with co-upcycled mixture-of-experts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:11.710677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:11.701774Z digest=sha256:4ebcfc300fc363d55b9d3db98e22abfb6e8668dd16a8791c456a199cb21f4110

Observation b92f1c4f-9973-4e2e-97ba-4a759a90d28b · outbound

This paper cites Hrvqa: A visual question answering benchmark for high-resolution aerial images.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Hrvqa: A visual question answering benchmark for high-resolution aerial images

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:11.280283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:11.808104Z digest=sha256:951351b7e6d87ed768f28bc2bd227fc3b4bb859b2ff2a21ed3eae3292553d4c0

Observation f22effdc-6b0a-420f-ab2e-5fbf73395dd8 · outbound

This paper cites When hindsight is not 20/20: Testing limits on reflective thinking in large language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement When hindsight is not 20/20: Testing limits on reflective thinking in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:10.860720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:11.883233Z digest=sha256:e9bfb5a9b897015b5de09ccea783effca403e292928820cd074402d330634d91

Observation e62d103a-67ea-41fc-bcea-d92ef42eb996 · outbound

This paper cites Mini-gemini: Mining the potential of multi-modality vision language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Mini-gemini: Mining the potential of multi-modality vision language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:10.502308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:11.956210Z digest=sha256:f0223873a9d8baa864f8250305d996659d66aa2a27737427514828eac96eeddc

Observation 8c015068-9151-440b-b307-0fba713fcd06 · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.979543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.979543Z digest=sha256:d407b5c39f9a36bdf0c8d3ddbeb8bb4842f42f71e1eb6eec8290c419b0d9d5d1

Observation c2d1b6cf-dd28-42b1-ad6a-2b21150e6076 · outbound

This paper cites Vila: On pre-training for visual language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Vila: On pre-training for visual language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.995784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.995784Z digest=sha256:4c8787e16148350251f939f2617ee9cbeff6bd2887894e6357ae03f6e943262b

Observation 91980365-2df3-49ee-a963-527143961fc3 · outbound

This paper cites Visual anchors are strong information aggregators for multimodal large language model.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Visual anchors are strong information aggregators for multimodal large language model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:10.222132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:12.012509Z digest=sha256:3c10cb68ba015183dd16e4e40ac678d0232db5bfeccd91ea01de38f3456202c1

Observation 93631198-2d5c-4fcf-beb6-697d249952ac · outbound

This paper cites Improved baselines with visual instruction tuning.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Improved baselines with visual instruction tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.037557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.037557Z digest=sha256:a2e09e04b7b9108afe5cc61ae8ca32fd6568d5082469093657ba731afab92f26

Observation 9f3c7515-4b67-40c4-9bf5-5d175824de0e · outbound

This paper cites Llava-next: improved reasoning, ocr, and world knowledge (2024).

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Llava-next: improved reasoning, ocr, and world knowledge (2024)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:10.003029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:12.062045Z digest=sha256:790b55e3755c1b2b38b6d1ef3f7dab4b384b38941529ceca3a30f1d74e3783a4

Observation 00e4b129-d5d7-472c-aba8-8778e2700f68 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.097104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.097104Z digest=sha256:33d848490fa6534e5ef2cebd3819838a9b9b24c3092bc4964afbc824ce1000cf

Observation afe2df36-233b-4b4b-8ef6-4474a7ff8dfb · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.120767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.120767Z digest=sha256:452afda8e78bc7bd9a7ebb21419bc6c61f9a3809cbd613475d2a4ee5a21fa31b

Observation 94f934a8-1323-4fa7-93f5-d281f039bb05 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.167451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.167451Z digest=sha256:e66e15b8731491e49d8b5a57c2c0cc21b68f43d1cc938b6db7a3723e5a353a51

Observation 9a24f198-bdbf-43fe-b694-f2d496b58f50 · outbound

This paper cites Task-to-instance prompt learning for vision-language models at test time.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Task-to-instance prompt learning for vision-language models at test time

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:09.758140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:12.227399Z digest=sha256:5fc0fe406dc058347f1b66d2722d50d69e447d5a86278a5ff6d09cac4c57839b

Observation c86508c6-2314-49f8-aa28-9a37f1f5abea · outbound

This paper cites TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.313744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.313744Z digest=sha256:a99e8eef6d5f2b165da2241b7c908f12b44b0994ba078a5d9bc1868b4f2b746f

Observation 2867ea04-cbad-4032-8c42-81fbca559218 · outbound

This paper cites Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:09.532909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:23.084530Z digest=sha256:9bfee218ea6556be160a986a7bb9ac418a908471753c893563fd1db46208935a

Observation abb0e6ed-2781-46e2-b6a9-a5564a20848d · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Self-refine: Iterative refinement with self-feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.398609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.398609Z digest=sha256:a25f1a14830dd839acb3475996c18d7ce93c150d0da54b16635089fbb5b6c6c2

Observation 6a1d21a7-d9c7-4166-b330-aaf67753a2ba · outbound

This paper cites Gpt-4v(ision) system card.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Gpt-4v(ision) system card

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:09.290415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:23.659485Z digest=sha256:b5abc886085c380573b3731c6e89ca5930ee25965f23abc99ded219433ae7331

Observation 65f80211-1a44-4583-b83f-d1d8bb65a67b · outbound

This paper cites Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:41:08.736476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:24.131642Z digest=sha256:fb685012f6dfd9d6111f548f3a648207c15c5540c733580c9f2a70eba9b7f051

Observation 4455c1e7-ce09-4d99-a68d-c76918055fbc · outbound

This paper cites Orienting of attention.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Orienting of attention

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.914294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:31.438152Z digest=sha256:c11fc029aea8940a2940c6ce41a7265b0bc6832ddb521b8b1b7f8cbe402f55b3

Observation 5b6a1295-b337-4ebf-b08c-5bc9c626abce · outbound

This paper cites Zoomer: Adaptive image focus optimization for black-box mllm.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Zoomer: Adaptive image focus optimization for black-box mllm

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:32.884746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:32.884746Z digest=sha256:321e6f39895c3fdc6e73541bcc3d38b07583aae42250572525e563a709811d02

Observation 5ae2e871-eccd-4eeb-95d6-0cdc5fddfa96 · outbound

This paper cites Self-Reflection in LLM Agents: Effects on Problem-Solving Performance.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Self-Reflection in LLM Agents: Effects on Problem-Solving Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:34.931637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:34.931637Z digest=sha256:260e2ab1010aafb95889a33f302cb01626e6bfa1f990f80474c558ae884556a8

Observation 74a5a5f2-1872-4cd6-9c98-002717c1c428 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.724614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.724614Z digest=sha256:0f5e25d557067961e29bcfbd0b45ba55a81f88d8ccd984dbca2cb0dd0accb795

Observation 4440d2ba-e57c-45c3-94bf-08cbf63fdc6d · outbound

This paper cites ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.885152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.885152Z digest=sha256:8ece424cb2cc9af0e907cb9a8b5c53448691e52099fb86c7f18cae83d2ab0803

Observation 02828600-11bf-4813-bdbd-0c7d797acf12 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.947540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.947540Z digest=sha256:b9270fff0517c96de1cf10a7e1e025248bf1c0107fd59ad6925ba93a0215385d

Observation 66162ec5-8ac7-4456-aadc-a8731e1eb92b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Gemini: A Family of Highly Capable Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.023592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.023592Z digest=sha256:fc9833011450b26a880160e0879dd27f36ec4b3e86eee487b6d83c844966db8d

Observation 0fbc42b6-179f-4d12-882c-11916fca7544 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.060732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.060732Z digest=sha256:9e8671a8f945fb386fcc5aed62169f78a2e8f059b60d744827f0e768908333c1

Observation f7416fb6-828b-4b77-929d-edacb27efd05 · outbound

This paper cites LLMs cannot find reasoning errors, but can correct them given the error location.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement LLMs cannot find reasoning errors, but can correct them given the error location

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.734979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:45.359104Z digest=sha256:3142ad9fcec6d94a2121aff5ef072b57ecfdf1e214081b3fa0df7da03b99577b

Observation 71067d3b-6c02-484e-9bcd-25fc548b3be8 · outbound

This paper cites Lever- aging visual tokens for extended text contexts in multi-modal learning.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Lever- aging visual tokens for extended text contexts in multi-modal learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.605872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:53.229790Z digest=sha256:4c4fe7e1a1eb5a7b2ce3d88acebff966871e959e8deaa4ecfae3fe06fce6bf66

Observation d9328b18-c20f-4bdd-9f52-cda4009c3eda · outbound

This paper cites The dark days are overcast: Iron-bearing clouds on HD 209458 b and WASP-43 b can explain low dayside albedos.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement The dark days are overcast: Iron-bearing clouds on HD 209458 b and WASP-43 b can explain low dayside albedos

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:54.884572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:54.884572Z digest=sha256:511489974c98b07391f1180a3475dacc5e0523c56e4bb2e6883f3b0b6f26f42c

Observation df6afcfc-3157-495e-9f19-70f11b2885d3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.542706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.542706Z digest=sha256:4ee355bb298a2a5202690ee0ed6a4604502f66ea15e4853224bcf6210a522bee

Observation e386970a-a070-4514-8af6-c3a44be88688 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Cogvlm: Visual expert for pretrained language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.871118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.871118Z digest=sha256:c230d99bebeb03cff91f58f028cd53750f799a5beeb1f97c4911ee2b8e2863a0

Observation a4ac4896-052e-41bf-9282-a694a9887d5a · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.434842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:55.914770Z digest=sha256:c7c6d3f172f3dc606defeffb090ebf1527ae12ddf4b9b373d31ea156ee3c65be

Observation 01d0dbda-89ac-45ec-a669-c6cc4003d3a5 · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.924464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.924464Z digest=sha256:e5671c2a38e8386a39523a88909da7df59c5913be4573e42dc47d454969e22d1

Observation e6b5bb8f-f7ee-4c4e-a119-32592b238f36 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement V*: Guided visual search as a core mechanism in multimodal llms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.302336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:55.953486Z digest=sha256:e418e3f76868f0de8bd62fee3247473b9f928380037d3b11f46a8194836b01ad

Observation 5220909d-b610-42d8-b132-db6725d818af · outbound

This paper cites Large language models can self-correct with key condition verification.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Large language models can self-correct with key condition verification

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.220091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:56.018805Z digest=sha256:c00a4c89363a71ef9a71099abe1e679627fc7a2a251c4ab1c99c724d7aea73ea

Observation c6cd0429-912b-4186-b656-91bc683d6c52 · outbound

This paper cites Graph- based unsupervised disentangled representation learning via multimodal large language models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Graph- based unsupervised disentangled representation learning via multimodal large language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.165808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:57.065875Z digest=sha256:53ab55731d4d01e8cd506515abf7d8500fb27d214d6ac1e4712540c8bfc5f9ae

Observation 208476fa-b365-4ac9-acdf-1f97ce4381b6 · outbound

This paper cites mPLUG-2: A modularized multi-modal foundation model across text, image and video.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement mPLUG-2: A modularized multi-modal foundation model across text, image and video

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.992755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:57.688044Z digest=sha256:dc2113b527b640d3235bc2b7f6e09f717c772255b83427af769dd81643d9458c

Observation 71d2c1b3-0784-455e-a21a-b735f55ceba6 · outbound

This paper cites Yi: Open foundation models by 01.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Yi: Open foundation models by 01

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.810469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:57.733682Z digest=sha256:cbf5ee8a252b7ce8196fe27abdc5b10350d75337da5ed5af263d9d6ffe277ef8

Observation 8f38b95d-d9c2-471f-b671-55de3805bb75 · outbound

This paper cites Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.775441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.775441Z digest=sha256:6245235b9d9bcecd5c8920fc5936c1121f34cb4307edf87620d39d21d8d48a9f

Observation 4db93dbb-ef84-4fd0-b971-c203f35a4d55 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.810777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.810777Z digest=sha256:dfb2904d3275e4f4df02c0dea5d3eb9e903c9430acf4c035fb7a444388f73814

Observation a6dccc8c-113e-4639-aa31-360a77cd65aa · outbound

This paper cites Beyond llava-hd: Diving into high-resolution large multimodal models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Beyond llava-hd: Diving into high-resolution large multimodal models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.656533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:57.861871Z digest=sha256:70bb56281c4b39b7de2e138a2c1abcca09591be042b058a49cffdbaff38397e7

Observation f0c9d9eb-6e6f-4cae-b789-a64b09108131 · outbound

This paper cites Wings: Learning multimodal llms without text-only forgetting.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Wings: Learning multimodal llms without text-only forgetting

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.543281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:57.905024Z digest=sha256:276c94aad3434f1aa79566482e73130d79aeeb69f265d92b68ffdecb002f0625

Observation 24a627c5-91d6-4e12-86c9-66262efea867 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.963146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.963146Z digest=sha256:5e9c6ff7e28daeb2d0b3001b2858401a53fb135934da3d0ecca742c7129657fc

Pith citing papers

Observation 9868b5db-7fe2-451f-b3ad-670da9d701f7 · inbound

Perceptual Flow Network for Visually Grounded Reasoning cites this paper.

Perceptual Flow Network for Visually Grounded Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.906515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:2cd27fef6d13b16207b5fe2d6a4428418ac4a023081bf59d2eedd241b3f6b458

Observation 21821802-1710-4230-af71-ed63661cdd3e · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.465339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:58:01.621488Z digest=sha256:4350f887c53198debc9b6576d45d631bf3c641ee5f469cc08a4a2cf1b6543337

Observation e1e983bc-ebd2-4237-8dfa-6156a944dcb8 · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.539600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:28:21.605646Z digest=sha256:917fbdfe84d00ac03b06fd8bba4b2f502b1d29a07dd80b734d29797fd867c83a

Observation 73bd0da4-65a6-45c9-8aea-a993e3567ffe · inbound

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning cites this paper.

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.983723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:38:01.819121Z digest=sha256:88bc40d576e4b081cfa3d0f6960361db7013edda0495ff50bef388fb1e4b076c

Observation ee861032-4467-4f75-a999-7ab0da331cb7 · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:22:37.728141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:ecffcb4420c3c5260cf2367698d253fbd1118ab7a63059a74a5d2a8236e49718

Observation 8c912615-7240-44c2-bdf9-0a1a1cb58bb8 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.340686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:6706dcbc911235ef25881ba707e5488b50cf27c1fcbfc65b5453024291dbd19a

Observation 6cdee9f2-199e-463c-a468-e0352d01435a · inbound

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context cites this paper.

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:21.312749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:01:49.904752Z digest=sha256:873e59fbcab04be6dab663e3273d3a1513a2a7ba6aa72bbede8066a685112b20

Observation ba875588-6208-4359-8884-f766a96d17ba · inbound

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning cites this paper.

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:58.472521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T13:24:17.538850Z digest=sha256:2c30d4da065eb7ff603247a00efe3daf0bb5b4d77183107a341578380fbe493c