Pith. sign in

Paper Citation Record · LEDGER

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

As of 17 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2507.10202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10202 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:40:40.290572Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:30:58.091365Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T14:27:24.213919Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 851cbd77-5559-4208-8134-c10a629ff47c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:34.116748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:34.116748Z digest=sha256:517694ba0531ed4fcf7cace13fa75afe349660a14b1f90da137458c3c3a2699b

Observation c9606cab-eb93-46b2-b7b9-3a150b16005e · outbound

This paper cites Lan- guage models are few-shot learners.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Lan- guage models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.177712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.177712Z digest=sha256:c032201e710145f73c50efa03cd4f21002263f2bd09509b7e36461b4a5651541

Observation e6b6550d-3f4b-4a44-b865-a89967a88b5e · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Emerg- ing properties in self-supervised vision transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.183319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.183319Z digest=sha256:9ada7eeb71f1e2ffa29d8ebeb3efb09b5384380bca7512d42fb2e387555cf8c3

Observation 9c6bd09c-3ad6-419c-84ac-49eb1fbe995f · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.767037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.189046Z digest=sha256:e286db1f0265860cdd3d88f041299c355744e2282fc6b8069137d892107738e0

Observation 5fcf4a7d-c972-46ca-980c-90db114f6b68 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.195059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.195059Z digest=sha256:e2a94f0ead2745cffc5e873c07015bd6fcc589e7d897c64d048ddbb12c15a318

Observation 62070ed8-9b1d-41c4-bde5-dcaef039002a · outbound

This paper cites Palm: Scaling language modeling with pathways.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Palm: Scaling language modeling with pathways

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.750712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.200279Z digest=sha256:fe39a10c06d74e51c4da83b9a42e81ab73eab828610dd2d77c783561067deb70

Observation 8519e9a4-b332-4971-bb49-87a39693aade · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images, 2024.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.732322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.205362Z digest=sha256:6a29e9ccaded0fef3552f33833e17fddf5b6a202ee0ef79016a4023a2dbe543a

Observation e48fe762-3355-4422-ad46-75249e10b127 · outbound

This paper cites GPT-4o System Card.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.210305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.210305Z digest=sha256:6b720cf1f60474254b16ffe84406035b7d51cfbf63dafda5db16a5b344e9b95f

Observation 499b9c21-a768-4a9a-ad3c-35bb8e1c4779 · outbound

This paper cites Screenspot-pro: Gui grounding for professional high- resolution computer use, 2025.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Screenspot-pro: Gui grounding for professional high- resolution computer use, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.714507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.215124Z digest=sha256:aa29b3ad18b2ec499985ed2b57b75f2a6f9721af3438c500f1842d58b0a1a79b

Observation c92f28e4-58b9-4c27-b487-cc5733b7bae1 · outbound

This paper cites Visual instruction tuning.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Visual instruction tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.697299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.219197Z digest=sha256:3dcf6b905d93d294b0aeadb81f217981d3bd1ce48e156d137076a084c47bc28a

Observation b02df333-5030-4d93-a675-c96a714e8558 · outbound

This paper cites InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.225281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.225281Z digest=sha256:2c87242fe78d858abc7af4a7ad866bbc67fd160d716d7ce49b1b8bd8cd187b07

Observation 04f92453-e9e4-4bf9-8fd9-dafccf4c7bdc · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.230169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.230169Z digest=sha256:a0661d4f2136767f35fbc3efa2ea71f9e73ab2534d550ec4bb28a67122d40325

Observation 66d49876-07ef-42ef-8be7-b7c6ef2ba733 · outbound

This paper cites INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.235754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.235754Z digest=sha256:fa11f1e1ebc7a95be620339ae95e3a3b16ba6a4cda04bc66fcf7085efd0cd082

Observation 683aaed2-2f3e-4711-80e6-a086e56633b6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Learning transferable visual models from natural language supervi- sion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.679550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.240963Z digest=sha256:47f23319ab4313d6f8158e04e56a00b65f335e27c7e9decc917c32f4e9b9489d

Observation bd70a414-2818-4b9c-a681-e26a7def89d4 · outbound

This paper cites Divide and conquer: High-resolution industrial anomaly de- tection via memory efficient tiled ensemble.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Divide and conquer: High-resolution industrial anomaly de- tection via memory efficient tiled ensemble

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.658842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.244930Z digest=sha256:22c4b35c6002987b6c2f4e8dcc916df2c33afe5c6e14beb8676d998c137be54a

Observation 127f03cd-3421-4825-a581-56623e3d8213 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.251610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.251610Z digest=sha256:d2c2f41502e116aa19bc8bf389609fccf55531dffea8405cfde7c3ab48b7323e

Observation 6a15456b-100f-4342-8ffe-776f179f7449 · outbound

This paper cites Fixing the train-test resolution discrepancy.Advances in neural information processing systems, 32, 2019.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Fixing the train-test resolution discrepancy.Advances in neural information processing systems, 32, 2019

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.643180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.256814Z digest=sha256:580b7d4ad7c528d715178d1292e2e6e951aaf087de72be4cf67a3a6b34ca1ad8

Observation 40b55d28-a80c-4212-86e1-64cb93d305b8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images LLaMA: Open and Efficient Foundation Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.261877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.261877Z digest=sha256:5dddc6dfbbb8396fde64378434c90e1d8d6f94d7bce220cf55f166c09eecf7f2

Observation 54d99456-7761-4120-999a-633afa533ff2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.266695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.266695Z digest=sha256:79159a031f23283b67496a6ae927c9f19fd3328650c143d9e2473893fed89068

Observation fd65673c-bd1a-4940-8b85-dd4bc318fbda · outbound

This paper cites Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.270764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.270764Z digest=sha256:43dd8d7cd18d54743e1ade819fecb2357897f55993f54d41dfff6bb02dc5399d

Observation 260882ba-87f2-4e7b-a439-4db25da6134c · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.276025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.276025Z digest=sha256:783340254f44755bea76d5310e34902e1940010a5bc6a238d08b3bd32f55316a

Observation fa3342fe-f379-4b64-8168-773eb6a7445d · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:40:40.626256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:40:40.280786Z digest=sha256:2d6a861c0326701f049eddae00bef2723979a0b9a8f46228da3d93807de299c5

Observation cde3fcfd-ef63-43be-8f42-583d70769997 · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Large Language Models Are Not Robust Multiple Choice Selectors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.285125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.285125Z digest=sha256:c638556a5ff1a7201a34c844eae3b21476a3392da5bdb861de8dc3b056bafffe

Observation 350fa9db-b1dc-409a-8387-d220a0c7aefb · outbound

This paper cites MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.290572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.290572Z digest=sha256:e0497e138e876999bb8357b2c2be55f29458005605fd1ab4f0b3d07f9d0290cc

Pith citing papers

Observation 4e58942a-f5e4-4b74-93c2-ab7e7c0e8702 · inbound

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models cites this paper.

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T20:30:58.091365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:30:58.091365Z digest=sha256:ac5722cae4cc397bc7b46025c3eb01db9ac1fdb0526b999032aea2896c767d9b

Observation 058f0829-90fb-46e2-aaf3-6ba1586c3262 · inbound

DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding cites this paper.

DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:27:24.215449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T14:24:48.938948Z digest=sha256:e93b650135b17af3f9968d7520ba837884b59d132e5efe34a58a3b58d0c97079