Pith. sign in

Paper Citation Record · LEDGER

LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2403.11703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.11703 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:13:03.169164Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:26.526436Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7dc71625-7470-4bd5-9191-21555ca9f18b · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 126

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T20:58:59.266253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:8c83a1199868d53868a906ba6e586720fa64727690b8f31176e2add8840213a4

Observation 4160050a-63f9-4da2-a747-14b2a3401ff5 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 158

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T10:46:28.791681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:a738e5f504e2d899b0b535dcd7d92c1afd3059b3428985ff4074a54f7ae105eb

Observation 7eb48ea3-aab4-46d7-89a3-843fd6901bcf · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T21:07:32.061335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:719ea6f4e279a2bd2994cc4219c63d8194e69ed8defe9e871b954fce20151b38

Observation 2a176405-3573-42a4-9762-43da5b515c7e · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:59:32.722067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:4df4c3a5f14cacb0b9be4b2ca661a19bd684ed58925b7bb55456d5a3f46dd2ec

Observation bff25321-8cc8-4f57-ac0e-fab80c0b6773 · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:12:14.716016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:35dc5ba4d931502ad768b90b826c3e65af8842228ff9735f8f1e046ac9ebc10d

Observation 4c9380e0-8f64-479a-9a71-12683b17e959 · inbound

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models cites this paper.

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T20:27:20.055623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:27:20.055623Z digest=sha256:0be21358dd722dbf7dac070f473117d55d15afa39c3bcb9cce659beabdf4e351

Observation 6aee29cb-5595-4836-be2c-d468dbc7c14c · inbound

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning cites this paper.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.873349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.873349Z digest=sha256:c92d0240cd8048cef6f3f3ff623b031e8ac79106830eb36eb3105c03d41ba38d

Observation b5ccce93-dd71-4f15-af6e-bde51b409499 · inbound

LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression cites this paper.

LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:00:24.637561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:00:24.637561Z digest=sha256:f82a075a0cb4889a836b8afbb2e3fd56affe3faad74917973ea7c79a67fc5822

Observation adf3e0c0-7b68-4965-945b-5221e8622735 · inbound

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression cites this paper.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.486326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.486326Z digest=sha256:d603fa89d2f758c4f27837860f21c0ebf1713ccb897782b0df628866e9156b50

Observation 243f9db7-0ee0-452f-875f-e68730c06957 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.921149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.921149Z digest=sha256:fa4bb424bfe5ff01d6ff9d47bbf44ff7eb0c0e769aacdb40a894d8bb1eb8892c

Observation e2fd5e40-6ffa-47b7-a945-2e14b621c775 · inbound

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI cites this paper.

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:16:27.913904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:16:27.913904Z digest=sha256:e27e98b18a842e6f4eeabe83f62d57920dbeb68f39f4f5fa2dc864f0a36fe13b

Observation ef32f889-edb0-484f-86f7-91a1467a4adb · inbound

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy cites this paper.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.872391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.872391Z digest=sha256:18f06fdc608266eb5f1dc0c6c0af46d3503c23b2a09670cda91ed3166b67a1f2

Observation 33b7389b-5513-4a1d-9481-c34d3bcb9e98 · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.866411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.866411Z digest=sha256:7a9e6c58811d265e5c0bedc2b7407e80a3013fafc76b97fd9f427e1e86782f2e

Observation 5647cc14-a07e-43cd-b4f1-8ea8d72e7d65 · inbound

PerLA: Perceptive 3D Language Assistant cites this paper.

PerLA: Perceptive 3D Language Assistant LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T05:52:39.162003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:52:39.162003Z digest=sha256:08ef145cc1e195d0e0dd1bed1297eb32b11d08893333e7fb870d94d78c9badfc

Observation d9beaaad-9d20-4101-9fe7-7fbb08b79fcd · inbound

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction cites this paper.

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:10.945578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:10.945578Z digest=sha256:5de3725959e0edcefa4f41947ae4c8be8fb2b9b5165d42cb88e42fd3f7fe1840

Observation 54a415b3-819d-41fb-977b-ce08a5b70f4b · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.547350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.547350Z digest=sha256:ce9bc012afde466f02bba3928098c6de62574841d4f17fba4e49158f234a3654

Observation a0472031-8831-4bed-a82c-bec1764fb727 · inbound

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining cites this paper.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.030128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.030128Z digest=sha256:11e052271429866c59a34b1d87e06bca45b5868aa63ef8841891ea5ece14023a

Observation dbb3d671-2012-410a-884e-ad8830cd6e8c · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.371099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.371099Z digest=sha256:e178ea0eda0bf1cc79ea7e18126f7aab844ec3edb3944bae1c0eaf87be661a97

Observation bd986c29-9567-499f-ab33-9bc948658b71 · inbound

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation cites this paper.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.309295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.309295Z digest=sha256:041323f3f1fe1846843da2791f2c61897ea6e89403c19a3d8e7dd8b802f57190

Observation ca19e749-1942-4d8e-87fb-4d3f901b9cce · inbound

MBQ: Modality-Balanced Quantization for Large Vision-Language Models cites this paper.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.300528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.300528Z digest=sha256:893ba07359a1575a7e45c50eeacb84af6cadda7b5fc6f0ba04310d052058ad80

Observation 6bfcd838-d87c-49f2-bbd7-cf1f88539ebb · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.060042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.060042Z digest=sha256:3d3d6878138315c4cf871135aff8b013d5f1a96d84272bb4bd19d1a5693ae28a

Observation 97733769-5075-41b0-9ffd-ae8eb34423ac · inbound

Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions cites this paper.

Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:28:28.962345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:28:28.962345Z digest=sha256:e0b6fd1c4f7435374fcc19f764eb5b84bd4360df42a74b57bcf7e3c70197e143

Observation 641d1f68-7aca-4ac3-b248-8dcc02586bb5 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.961445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.961445Z digest=sha256:a26b9e737af0c2418f5d697605f1e2d01d1d01bda45ed1f5c1cd96e10a8840c4

Observation c73ca218-7d4b-46d1-83c3-fb6010d0fa22 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.962910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.962910Z digest=sha256:968c90ee110302ff8d29a934309c772f7368e1482b75ba8d2836cde745550dad

Observation 5b74ac38-e830-468f-a941-ae1a69344cba · inbound

Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks cites this paper.

Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T01:01:03.854216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:01:03.854216Z digest=sha256:5570a744f63fc8ad2c7a3e8eba911125267ca62cad5c8e2eacd72e1a5ef0285d

Observation 4207f9a5-e042-4533-96ab-33b5cfcdf685 · inbound

RESAnything: Attribute Prompting for Arbitrary Referring Segmentation cites this paper.

RESAnything: Attribute Prompting for Arbitrary Referring Segmentation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:13:03.169164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:13:03.169164Z digest=sha256:4ff5ca52b8211cf85f41e8011ae998f03249cb742e1081303d77f9b84fe91251

Observation 2a7fd5c2-e049-4bdb-8a1d-5184c95a36b1 · inbound

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs cites this paper.

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:35:13.220175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T05:35:13.118221Z digest=sha256:940e246aa5093860e0c5d8e7843c6eeeb36552e48f53727deace41a310c5f555

Observation a948cb9a-ed4a-49d7-a9c0-9c532f81ba55 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.624340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.624340Z digest=sha256:4d8fe4699a57d82544b7ccc038e993e0ed01a615cdb3c404d0a23ceb43f2163b

Observation d30a2810-ca8f-4506-841d-cab57a499d09 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.966497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.966497Z digest=sha256:54c1cbe160a546400a696cf41e397b71cec9dd5206d88882a3954e5908f3cd72

Observation 04f360b7-f5ba-4a69-b6cf-8a1077010caa · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.114851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.114851Z digest=sha256:bebbafd1b09ed21dd3906d23e29ee13893a8155995cf5479a394df74d267b97c

Observation 8a51f921-683d-4c95-aaa1-d22e69acf980 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.946036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.946036Z digest=sha256:4d11468b15dce6a4f38241c192d25625159654c0866041b5325edf43bf1949ac

Observation c856420f-0375-4530-bb92-9893575ac663 · inbound

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions cites this paper.

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:46.494815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:46.494815Z digest=sha256:5a86984df3a3b5b5d89bee4df9ee915d0bbbaa0a405dc444fe004aaa90c8f02f

Observation 1e2f5763-d0fa-42bd-93e6-b8142f783d72 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:33.650037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:33.650037Z digest=sha256:95ae9d47b180ab7f3a7d35168954ca364e5343fab8644bc59dda1490ac99498c

Observation ee24e842-e62e-47ed-ac14-e8fe6817bf62 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.926907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:40280ba3a5e1cbe5c7b68a5ae62298a2b859537d0b4b3589cd090d72d1c71e48

Observation 1877bc03-d0cb-490a-927b-beb65ecab6fa · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T09:40:35.631188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:40:35.631188Z digest=sha256:9c38a5bf0e5487c7309f3c583c98d4f3c13b98cd9df6d9ebf14c6fe239fc5ca9

Observation 653deced-ed06-4e0f-b05e-96b9ec49eb7f · inbound

Less Detail, Better Answers: Degradation-Driven Prompting for VQA cites this paper.

Less Detail, Better Answers: Degradation-Driven Prompting for VQA LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:05:47.957190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T20:17:01.867903Z digest=sha256:91020c87fb1ed9702eefca13ae34fb4e5959db005eefc56c1115dbc142602e50

Observation 7d1d40ee-1341-44cb-abad-283d687a1dc1 · inbound

How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A cites this paper.

How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:53:49.169616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T22:52:28.333209Z digest=sha256:d648f595e7921fa670890be20e0db57e3a01271a1fecaa6f6894d3a54a0c0b84

Observation 9de7d28e-7dd8-4052-8e29-2b5b9566c68d · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 209

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.689687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:e044f26d72268042d8337a7170b1316220d0c54e209a4ecc642190be77f068db

Observation 08aafe95-02e6-4d43-93a5-4bcb90d34c15 · inbound

Self-Prophetic Decoding to Unlock Visual Search in LVLMs cites this paper.

Self-Prophetic Decoding to Unlock Visual Search in LVLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.895231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T12:53:40.783281Z digest=sha256:2bd3063b870c28a604f1c50e02cd98814e7f216a53073c161561598554c56cfe

Observation e7bdc8b3-0d3d-42e5-bd3d-270ddb991127 · inbound

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics cites this paper.

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:26.528104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T10:51:38.455604Z digest=sha256:d822d07ec2487fbdaa03f68cdbd93079baeab70eb728c5a5e9c7503ce12be148

Observation d0d2ec27-054d-4647-b706-b20d7448e68f · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 228

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:7dd68234047cd8acf49ebdf284c32632ba44df08c6a0a1c93d1e77d2ae8ace7a

Observation c955fe02-0de6-4385-83ac-b075f505bd63 · inbound

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning cites this paper.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.341288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.341288Z digest=sha256:4872b35d5060a45303d53f31c6f50608d716bbb9ee16f764925b5db9be55765f

Observation ba25bc3d-d543-41ea-ba6b-fe2f83f6904b · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:58.293552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:58.293552Z digest=sha256:9e9c0a68adf63e7527856d3205d16f0546febe1733f61f757c00e5380af5012c