Pith. sign in

Paper Citation Record · LEDGER

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2507.10202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10202 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:40:40.290572Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:30:58.091365Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T14:27:24.213919Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 851cbd77-5559-4208-8134-c10a629ff47c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:34.116748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:34.116748Z digest=sha256:04b6271336260a3b4aa45ffd52b3c538bbc6a0e9e90fbda52f14d9065b1d5cea

Observation c9606cab-eb93-46b2-b7b9-3a150b16005e · outbound

This paper cites Lan- guage models are few-shot learners.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Lan- guage models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.177712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.177712Z digest=sha256:77d403d322582222f37b649601c4dfecf4515abfe544099eefa92011534ff831

Observation e6b6550d-3f4b-4a44-b865-a89967a88b5e · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Emerg- ing properties in self-supervised vision transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.183319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.183319Z digest=sha256:80e32a9f043ae5474b996ce4d2df3151cf55d6cb7555091c44b4d41f4c485081

Observation 9c6bd09c-3ad6-419c-84ac-49eb1fbe995f · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.767037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.189046Z digest=sha256:824d5890e47c62f3cd476296f251447f6794f0f01a60aa0bc99178925cedd855

Observation 5fcf4a7d-c972-46ca-980c-90db114f6b68 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.195059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.195059Z digest=sha256:16643c474d2c2711956af2a7552e2513c7a88935b7e422f5c8903119a7a035c3

Observation 62070ed8-9b1d-41c4-bde5-dcaef039002a · outbound

This paper cites Palm: Scaling language modeling with pathways.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Palm: Scaling language modeling with pathways

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.750712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.200279Z digest=sha256:9d3151c31907dae8af68017c21562524197aa98dac84b39b2fc9918a065bf015

Observation 8519e9a4-b332-4971-bb49-87a39693aade · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images, 2024.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.732322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.205362Z digest=sha256:cb068dd7a8a93bb3dec633ce0d35a3a555caec92f0c3911a51cfd27240a27188

Observation e48fe762-3355-4422-ad46-75249e10b127 · outbound

This paper cites GPT-4o System Card.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.210305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.210305Z digest=sha256:ee541f4a54db462a4862ebacc879da08030e71d86929b6ec80674ce19704476c

Observation 499b9c21-a768-4a9a-ad3c-35bb8e1c4779 · outbound

This paper cites Screenspot-pro: Gui grounding for professional high- resolution computer use, 2025.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Screenspot-pro: Gui grounding for professional high- resolution computer use, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.714507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.215124Z digest=sha256:4f1ee6c91b82ee9d30e5d3e7791775b0d993dc9eed49012133d15aa7155cc79d

Observation c92f28e4-58b9-4c27-b487-cc5733b7bae1 · outbound

This paper cites Visual instruction tuning.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Visual instruction tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.697299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.219197Z digest=sha256:9d0dba79e659f67d679d01ab018d02131595500b07ee3d56a924d3c3db8827ea

Observation b02df333-5030-4d93-a675-c96a714e8558 · outbound

This paper cites InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.225281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.225281Z digest=sha256:e0b33671cdaf4ac06038af97c0ada56abd5cae67b4560d7c96ba10ba7d9e2fbc

Observation 04f92453-e9e4-4bf9-8fd9-dafccf4c7bdc · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.230169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.230169Z digest=sha256:35dd766879c48ff3208067e025fcafa2b61e0c802327532353b99b7d885df525

Observation 66d49876-07ef-42ef-8be7-b7c6ef2ba733 · outbound

This paper cites INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.235754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.235754Z digest=sha256:3a51738dcd620578dcbce122860fbc489b565b382f18c9a2072757e3eba04e97

Observation 683aaed2-2f3e-4711-80e6-a086e56633b6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Learning transferable visual models from natural language supervi- sion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.679550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.240963Z digest=sha256:66b887de8e0c45e2ad6a56a298b2811f54a31bd8c489f7c09de90ef06c66db95

Observation bd70a414-2818-4b9c-a681-e26a7def89d4 · outbound

This paper cites Divide and conquer: High-resolution industrial anomaly de- tection via memory efficient tiled ensemble.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Divide and conquer: High-resolution industrial anomaly de- tection via memory efficient tiled ensemble

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.658842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.244930Z digest=sha256:3283e95296d6d781df71eeeda5766c481e18ace23352f9542513aac920f919e0

Observation 127f03cd-3421-4825-a581-56623e3d8213 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.251610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.251610Z digest=sha256:fb64590241efce248c9844890c940195d6b9dbfccd675e235d7b27dbf00ebfe1

Observation 6a15456b-100f-4342-8ffe-776f179f7449 · outbound

This paper cites Fixing the train-test resolution discrepancy.Advances in neural information processing systems, 32, 2019.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Fixing the train-test resolution discrepancy.Advances in neural information processing systems, 32, 2019

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.643180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.256814Z digest=sha256:32266a702e79386698ab54ab77411e8e69a815d2d02c807eee274d8640b68877

Observation 40b55d28-a80c-4212-86e1-64cb93d305b8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images LLaMA: Open and Efficient Foundation Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.261877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.261877Z digest=sha256:3e2234c3b6c8cf510c22164b7a665ea4c78393d9505e6b39fad216b1e82853cc

Observation 54d99456-7761-4120-999a-633afa533ff2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.266695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.266695Z digest=sha256:29e3738636a5a495eb4c8616673407db090e229388dd632f269cced7fa6dc4e1

Observation fd65673c-bd1a-4940-8b85-dd4bc318fbda · outbound

This paper cites Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.270764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.270764Z digest=sha256:fbe07d0db2e5694733d9509c0dc48b076d1a88bc8cc90599f055866c80cd01ed

Observation 260882ba-87f2-4e7b-a439-4db25da6134c · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.276025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.276025Z digest=sha256:eb7b6106f22067131a6505f7c6ab93adcc426ee4f9cb68bcef8f98cb6d672875

Observation fa3342fe-f379-4b64-8168-773eb6a7445d · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.626256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:40:40.280786Z digest=sha256:3ee28b6cc77627ed5f9b702ad21d2441bf0bda8d235419b30d98219383984273

Observation cde3fcfd-ef63-43be-8f42-583d70769997 · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Large Language Models Are Not Robust Multiple Choice Selectors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.285125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.285125Z digest=sha256:f536f047d17e343840684cf186056a5287607caf03d0c008a723649f06591bba

Observation 350fa9db-b1dc-409a-8387-d220a0c7aefb · outbound

This paper cites MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.290572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.290572Z digest=sha256:a63b5e4d8d4f35b8341885f8b562c59a7d1fb480bfdfd27345de3dd21bf3c9aa

Pith citing papers

Observation 4e58942a-f5e4-4b74-93c2-ab7e7c0e8702 · inbound

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models cites this paper.

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T20:30:58.091365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:30:58.091365Z digest=sha256:0cd8c3f59d280c5e28ee95e6744c07798ab0a8e718e10e755d0c14798f7d1146

Observation 058f0829-90fb-46e2-aaf3-6ba1586c3262 · inbound

DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding cites this paper.

DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:27:24.215449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T14:24:48.938948Z digest=sha256:9442f8e5dba969e1584fba2e359efd5caab775e823aff11e46a007142fcbb6a3