Pith. sign in

Paper Citation Record · LEDGER

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

As of 16 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 9 inbound Pith citation observations for arXiv:2412.13871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13871 v2

Coverage vector

measured 100 of 119 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:46:59.931076Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:39:35.380728Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:43:00.317941Z

Reference resolution

100 of 119 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved91
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d938658-7d66-4962-baa3-f104b23cdbad · outbound

This paper cites https://huggingface.co/datasets/xai-org/RealworldQA.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer https://huggingface.co/datasets/xai-org/RealworldQA

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.336510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.336510Z digest=sha256:d34a90eb8bd44ee5ec3b28647e2c8650c49f6acd65cc5ded9a69fce9b4bd96e4

Observation 8726ff5f-9218-4400-89a3-342f28550a86 · outbound

This paper cites https://openai.com/blog/chatgpt.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer https://openai.com/blog/chatgpt

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.342422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.342422Z digest=sha256:fb731811dca3c48e3a7dd8bf5ba856e34eb1ea3ebd534649beefe3301d76cff9

Observation 89178c1f-e3ba-487b-a310-eaf481564107 · outbound

This paper cites https://huggingface.co/datasets/laion/gpt4v-dataset, 2023.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer https://huggingface.co/datasets/laion/gpt4v-dataset, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.347956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.347956Z digest=sha256:de1e27cebbeabe34c04b4cda6799f55ad8df8ec6e245b2e521ac13ec8673dddd

Observation cebf44c4-247f-4149-a860-7f10a008f760 · outbound

This paper cites GPT-4 Technical Report.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.353376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.353376Z digest=sha256:c96deaac5e8902c8f0f3cee9fc8cb57c0455a2b0331cd014334924b4e8163e2e

Observation a20896bd-eb2c-4f37-b4a3-c15722d4b43e · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Understanding intermediate layers using linear classifier probes

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.359870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.359870Z digest=sha256:f46e9aef0880cf95a35e81b6829a542a0cd93cda593916f3d534d19f1e2c7cf5

Observation 4f9bdd3b-9aa3-488c-91f3-2bd087fd22d3 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.367120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.367120Z digest=sha256:48208170bc9099120d2ceeec1bcda7248607771bd44254bfeadfb8cf5d9bc29a

Observation e26f5bef-9fa9-437d-baad-921084d76ef2 · outbound

This paper cites Multi-label cluster discrimi- nation for visual representation learning.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Multi-label cluster discrimi- nation for visual representation learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.373279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.373279Z digest=sha256:0064eb8a86322ddc04b6b3b76a5d1f0cd12b657908703902dbc2a1f128e3ae92

Observation 328a96cf-04c3-4bc5-9b45-c36c00a4d7d5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Gemini: A Family of Highly Capable Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.378478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.378478Z digest=sha256:c69d56a76885153c0a88e3c5e901ec47fab325e1b6ca9ce3e2abeca11f6ea45f

Observation 12af006f-8cf2-4386-84e3-c77c47027708 · outbound

This paper cites VQA: Visual question answering.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer VQA: Visual question answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.385041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.385041Z digest=sha256:7dbbe4816f0f08ae813984480ce5880243b38775c71456eadd426a513e7f2ac2

Observation 440a624d-a99e-476b-838c-c8cbc026f9c8 · outbound

This paper cites Vision transformer for fast and efficient scene text recognition, 2021.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Vision transformer for fast and efficient scene text recognition, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.390439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.390439Z digest=sha256:a118bfe45ff7ead5e22bfc981c58132f79b4688089d9e018f5256aef4d0e04b0

Observation a81f5960-63ae-4a89-808e-bfbd42309127 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.395415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.395415Z digest=sha256:b6a4dc5534f60c0c539b9a9e9039d617ec1d8aa93b911c7b313d8e82cb4f583b

Observation a03e2fd5-77dd-401e-82ef-045ff8e0ce75 · outbound

This paper cites Introducing our multimodal models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Introducing our multimodal models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.401039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.401039Z digest=sha256:789e9cc7f2f80e7021b5ab2e597e59b10d6f04bb35ef53e8dcaa2a4eeef4e855

Observation 2840d8cd-8db7-4540-9e21-53c2ad50dee7 · outbound

This paper cites Token Merging: Your ViT But Faster.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Token Merging: Your ViT But Faster

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.405865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.405865Z digest=sha256:1482e332e131ab76b911e5efdb521447cffae166353acc7c17341216c76382ca

Observation 9beabfef-632d-4ce8-ab61-ace1893a7885 · outbound

This paper cites Burt and Edward H.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Burt and Edward H

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.411828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.411828Z digest=sha256:bc4bdfa592b5504e71404ffa84fa77e7247ad91e081ff46923798088b5fa53b2

Observation 9e1bfaab-7b61-43a5-b432-6ed8d8abb769 · outbound

This paper cites MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.417347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.417347Z digest=sha256:b9e4b62701c76caf81fff61177e5d6926bf96f4165d1428edf1c29d390e06098

Observation b18ae7ab-6da3-4f44-b5c9-9058b371e453 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Honeybee: Locality-enhanced projector for multimodal llm

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.422706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.422706Z digest=sha256:32535895437e2d387ad86913aeb007e54c17e06a587c9768074303e62629abb5

Observation 40ecb40d-8b52-40eb-8f63-cca626689cc3 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.427857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.427857Z digest=sha256:2204bf98ec2a9eb4594cecf6aecddd3dbf2e5dbba96283b21eaa35caf7696f60

Observation 0b2efd09-6a65-4065-bb2f-bb74d5075251 · outbound

This paper cites Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.433530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.433530Z digest=sha256:181b05032eabd071134f74ae89e2311c73e47ef73d6390b04592fce6bcc3346e

Observation af2d6da7-065f-4bad-873b-d9af5e5e3b80 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.438565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.438565Z digest=sha256:f4a8f86c35f4f3f4bd2a99659313e6ec0b4ef2c5819e50bc42adc3dafc7fa4dd

Observation 40f7086e-4021-4ef2-9a6d-35cf011e9801 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.443993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.443993Z digest=sha256:64d50381cbf0d95c82e797f4cd50d6e18f3f1808cb9671bd8037dd4feeb3f745

Observation 6400ae54-8e11-4b0e-974a-666d9771c0b1 · outbound

This paper cites Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.449186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.449186Z digest=sha256:8ba848a5ad51333355a578b91df394e2c5208ae712fcb94e9c818ef66669fa99

Observation 6d27094e-ac80-4478-9cc7-737b3866e3d9 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.454414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.454414Z digest=sha256:8231ec0ca6b459228de32a2755f33269c9bbf91e9c5865d8aed19db7f167019d

Observation 73e02ea4-56cd-4aeb-be70-1336a6b14ac4 · outbound

This paper cites Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.459532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.459532Z digest=sha256:4d528ba257abb374b76668999b3849fc590ae7a163864df85f524d014f3cda39

Observation 81438696-186d-4102-b4a0-5c48293d3d70 · outbound

This paper cites Xtuner: A toolkit for efficiently fine-tuning llm.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Xtuner: A toolkit for efficiently fine-tuning llm

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.464426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.464426Z digest=sha256:9b084834c6262225bdd1872ca44b3b5dc7bbab07227433fdebd6e46e7074cc07

Observation 33b02e12-757b-45f5-823f-fd29447ac61a · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.469635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.469635Z digest=sha256:82a0f3924e6e7f4fb9aadccd6c77c7be7e0a74f6001576816fd50f081433ca76

Observation eb885f6b-a024-449c-a5b2-f291cf5bd7e3 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.475234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.475234Z digest=sha256:c5e3be049a198dc13fba5379b83adc3aa8b81982fcecdd356d38536b426fc04c

Observation f5e57181-18d6-410e-af4b-e35b9ffeb216 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer An image is worth 16x16 words: Transformers for image recognition at scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.480439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.480439Z digest=sha256:613a146b895763b6b2d98921249a523d4174aa8f1466d217d9698922b76baec7

Observation 4b8e2d69-78ca-461b-93c9-914821742653 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.485449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.485449Z digest=sha256:1838a7ec2120391a093badff354b62c0d17fb1d22424b72de3b635c2804220f6

Observation b3c72ca9-ce90-4dfd-b677-93297f6c9d68 · outbound

This paper cites FeatUp: A Model-Agnostic Framework for Features at Any Resolution.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.490711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.490711Z digest=sha256:5f56c9065b6e548ccf4fb5ffe21698276ab95f4c38908504fe274e900b6fc211

Observation e369f5a0-d97b-49f2-b739-e8b58a348070 · outbound

This paper cites Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.495874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.495874Z digest=sha256:67b19afd9aa5ca917c145e4ffa5a66c39a91fe60e92a88d15fd37999b2f1764e

Observation 5f912bcb-b483-4bd9-8eef-d73b36f1e3da · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.501091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.501091Z digest=sha256:96c8afe1decefd82a6eab6029ef2851790ff2d700652418378d0b17ad011f2ab

Observation 6000433c-cf29-4f0a-9898-d7fc1ee58c29 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.506302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.506302Z digest=sha256:09d7df2c914804624c968adb184a224600374802e817f2f57746e62bc1be92ba

Observation d242c4f6-fc85-4337-ab09-75c0a09799d2 · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.512054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.512054Z digest=sha256:819bec403d9d9fdddf5609b2fdfed0d649f814c3cae5e60bdc386499e7f6603a

Observation 3bab988e-6a6a-4d8c-88ca-92e65352d052 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.517122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.517122Z digest=sha256:6e745b6b4efc54f884145cd46b0cebfa4c2da0144d8855df5d8079306ba8a869

Observation a978c192-42c0-42ee-9b53-13584642a8fc · outbound

This paper cites LLaV A-UHD: an lmm perceiving any aspect ratio and high-resolution images.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer LLaV A-UHD: an lmm perceiving any aspect ratio and high-resolution images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.522099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.522099Z digest=sha256:f70fdbb7b19faf40cc95e0ecd53052381d45612701b7a7a538e27faeb26bc3be

Observation 464e57b3-2e21-4c87-a474-64c4a3faba6f · outbound

This paper cites VizWiz grand challenge: Answering visual questions from blind people.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer VizWiz grand challenge: Answering visual questions from blind people

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.527115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.527115Z digest=sha256:a7eb1adb32fa04d4a63d1e04e9fa797c9112bcf2444503f950c6457d56840938

Observation 80a326f9-93c0-4c86-a0d9-db68247c9504 · outbound

This paper cites It Is Likely That Your Loss Should be a Likelihood.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer It Is Likely That Your Loss Should be a Likelihood

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.532147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.532147Z digest=sha256:af4b4c4763b9884a3d810a44c57cb42f67eeb90243e0272ce6869d08c12a92e8

Observation 57fa13cd-02f6-4dcb-b617-ccf4beada08f · outbound

This paper cites an unresolved cited work.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.538878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.538878Z digest=sha256:95d1122d272efbab5fcc158d606877ad530eef1d403f86f6c7fdd6c26c603249

Observation 4534fd55-d7ee-4fe3-ba40-675990ad0658 · outbound

This paper cites Deep residual learning for image recognition.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Deep residual learning for image recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.544967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.544967Z digest=sha256:7e5ac44b93b06b89ce860c64a2803b9d86c58cca1f997b12880386e38e4a89dc

Observation d745fad4-e8ee-4ec4-bc30-0a3c9bfed479 · outbound

This paper cites Mask R-CNN.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Mask R-CNN

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.549807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.549807Z digest=sha256:3734b6f96206b92894edf168704a0b6c7dc4ecdff8fd23c038cf257862811e6b

Observation 6c829448-aa43-4727-aa99-45ef772a8d61 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer CogAgent: A Visual Language Model for GUI Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.555350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.555350Z digest=sha256:b302dbc6dd6cb34ad7fda69f44fef5c7ff75e498cdfb76ce686077178817b4b4

Observation 3b99463a-c2f5-481b-9225-59fba3f95ded · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.561691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.561691Z digest=sha256:2e0740eb07787e67d1095e3cc713208eb6cbd78ce39f9a7fb33d9554cbc11644

Observation 9610f394-12d8-4f88-8063-18e84da21ee6 · outbound

This paper cites Synthetic data and artificial neural networks for natural scene text recognition.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Synthetic data and artificial neural networks for natural scene text recognition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.566778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.566778Z digest=sha256:ed413bd1455c91e3a236d27f08192a3742115b500ba2a5c740a6032e6a7b9902

Observation 53fefc95-47de-4e49-ad12-57cc8dfd8bbe · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.571912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.571912Z digest=sha256:5a27f3a35d331f7390b1cfcdd95ed1e0a0e96b17e283cf06b3278a14fdd993d4

Observation 44ab31ee-d5f2-4eb2-ac77-ed06829912ac · outbound

This paper cites DVQA: Understanding data visualizations via question answering.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer DVQA: Understanding data visualizations via question answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.577746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.577746Z digest=sha256:bdea673bfa69a9946d1205918f9ded9a3657d4d573327bbd1147e32616d2d0e8

Observation e61f89e6-39e6-49d9-ab70-abf183630820 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Referitgame: Referring to objects in photographs of natural scenes

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.583374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.583374Z digest=sha256:88a18b05a27c33a91a230b4be4fccfc9b05f2e97f9235aa4bd9b37b5458be233

Observation bf0c10f1-b32e-4690-9f68-1db2db626a10 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer A Diagram Is Worth A Dozen Images

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.589023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.589023Z digest=sha256:06cbd6c4c4fd72f67cd8ac15cc31557f9b6aa0cbec6d3dbe31e48a61e669d632

Observation 59a84b43-14ec-446c-8aa3-1ed65a65ff0b · outbound

This paper cites A diagram is worth a dozen images.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer A diagram is worth a dozen images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.595507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.595507Z digest=sha256:31d7e1674df4f8273fcfd2fd93f7983f638bdbf445c5cf32b0c3fcccb3c02423

Observation 844a8005-5df7-4d1f-a926-ccf5b365c54e · outbound

This paper cites On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.601941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.601941Z digest=sha256:3cae8a3e51b4258f863b110c70bbcb43738b8cfc0be10a59055f58d5d561f2c3

Observation 0d4c8b8b-f4a1-41b8-9c54-51e05732990e · outbound

This paper cites Ocr-free document understanding transformer.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Ocr-free document understanding transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.608473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.608473Z digest=sha256:9b95068119339ffcfdefd396de3a4a2621f683e5e38516958efe3fab06ff8457

Observation ccebcbdb-3519-4b9c-8dae-d85385937726 · outbound

This paper cites Segment Anything.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Segment Anything

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.614445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.614445Z digest=sha256:4c400e640c22ca4362fc01fb15f1f294e4da972d7442c628bcade0a4cb039d63

Observation a3300d09-90af-478b-a43a-7b6499cde8d3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.620717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.620717Z digest=sha256:411c1347d3ec607a8bf4fde1bb49b2cc113aa6e912c64e1e0e1475aa2b370c0f

Observation 2c7adf03-c513-4f38-89c0-8eaf96a2b13f · outbound

This paper cites Pix2struct: Screenshot parsing as pretraining for visual language understanding.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Pix2struct: Screenshot parsing as pretraining for visual language understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.627051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.627051Z digest=sha256:ecb03453bfc9a095e8dce09ab1a8ca971880f9acc4686317f03a2d442bf3b076

Observation ebdcab75-016e-4367-8fde-6652120f67dd · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.632684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.632684Z digest=sha256:b34facd21002f49b29672dd755217cd4b05a9f77ab802e629934d1dd5ec99650

Observation 95274e59-dab2-45e1-8daa-0e440bc6809e · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer OtterHD: A High-Resolution Multi-modality Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.639369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.639369Z digest=sha256:09d82128991f6920b2fcc590d42210091fcbd2f178295855af2890e29c5fc7ab

Observation 2570af2f-4935-46e7-bcce-0fcfa510b934 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.645252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.645252Z digest=sha256:b08786bf86ea58728b767b2e33d954d050c4b049f3e46c453e73b0a2043eac29

Observation fb05bf86-a2c6-460a-8300-fb344eaa035b · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.651832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.651832Z digest=sha256:a40fcdc7938ffd18d75b044c5f20638c26eda6dff8243a232564ed3ee9241c4b

Observation b3c1b57d-9700-4cc1-bf89-c0f4903b85ab · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.657618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.657618Z digest=sha256:a5143105e7dcef2577675673409416f41d3f01572ee30b9378ef74224fd6b4a3

Observation f99a26e0-6a84-471e-9778-0f011d2d64ae · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.662855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.662855Z digest=sha256:c78bcd8adf78793b9677260b836ad8216115dfe4d9797378ded6399c6b246108

Observation fabb162a-e7d9-4bd3-a0f6-dc13ef9390b1 · outbound

This paper cites Vila: On pre-training for visual language models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Vila: On pre-training for visual language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.668249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.668249Z digest=sha256:7401b6d160c0a22be2c32ad4f85bc7cad03929cf3ade8fa2ff7279d399258782

Observation 9a0cb0c4-3b8b-4c2e-ae79-f1ebbd2aa2f7 · outbound

This paper cites Microsoft coco: Common objects in context.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Microsoft coco: Common objects in context

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.679796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.679796Z digest=sha256:5374724a85fd05ba7d19bb4e2aa4dc0a2ce6799c3800b3afff69ee0033ac0891

Observation f17a4449-0be6-4843-b37d-470c68708fc9 · outbound

This paper cites Feature pyramid networks for object detection.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Feature pyramid networks for object detection

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.685247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.685247Z digest=sha256:3c9e4b006c47b207a96ae261754a43aa068ffee5489a9f3ff1dea862a512188b

Observation a28a15d7-6341-41e9-9f43-9fd5a5a1b2c0 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.690390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.690390Z digest=sha256:1f7af13f93b521601bee95c7dc47959c51eeea346e53260c2f13c3a57ce10444

Observation a07460a7-c114-436b-b1ff-0d8207bc7a00 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.697300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.697300Z digest=sha256:678d260f3b0f43e6212b7e4d6ae57d733b3cb823a4d634151c39c3adef9c4621

Observation 4a415545-e2e1-4593-8ba9-e0cca7385aa9 · outbound

This paper cites LLaV A-NeXT: Improved reasoning, ocr, and world knowledge.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer LLaV A-NeXT: Improved reasoning, ocr, and world knowledge

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.703540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.703540Z digest=sha256:1723873eb14ff6d5944f1a181114ecf69488bf63e6424fb52091f37b63115fd4

Observation f6964874-820b-4799-b1a9-182c9f0838a8 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Improved Baselines with Visual Instruction Tuning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.709359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.709359Z digest=sha256:36f8d016cf334cdc95b85a9570e51b3faa8da05dac24bda5ada245749c179fa6

Observation 130ca913-574d-468a-9e4b-3dcc5753d830 · outbound

This paper cites Visual instruction tuning.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Visual instruction tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.715679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.715679Z digest=sha256:8eb9ee9b6504b918e84d19e36fec0a10fec4f6ff6ddc7dd92859ee1d9c2e1bb4

Observation adef2197-0dab-46d3-a36c-e5e5a1026e85 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MMBench: Is Your Multi-modal Model an All-around Player?

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.721647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.721647Z digest=sha256:a367ab3a65d7448c487b5458c1ef4887028dba6116b65a67d85b1496e8af61e0

Observation 700232e4-a985-43e0-9b14-21cae2eeb0d8 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.728554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.728554Z digest=sha256:c3f212927ef073ea8923e385c8099a261f031b3e07aba81571e581d891363643

Observation b9c8882b-a001-4348-bb7c-b9d814c12715 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.735908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.735908Z digest=sha256:873845d775348a840dd448f6e72652b363c614d547784929fc9a75ff5fdfc994

Observation 570199d5-1f0c-4c73-b2e3-9974e24a9c9b · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Swin transformer: Hierarchical vision transformer using shifted windows

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.741534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.741534Z digest=sha256:52c96ac4e54180e7183248525aff35e5ccc7fd695355fcbcf2035d37b68b310c

Observation d7927ca1-ff8f-4ae2-96ca-87fe29363896 · outbound

This paper cites A convnet for the 2020s.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer A convnet for the 2020s

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.746975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.746975Z digest=sha256:cf7b76cd837e2403eae0309f9fa48beba1668de31e4be29c9034a874e2a9501b

Observation b9d2e19d-f8bf-4f2d-b833-61b343c53b28 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.752481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.752481Z digest=sha256:7f01448a33d930b0951371fab9e41ba84d715391115b7adead6d36105582c6fd

Observation 62e9c1e5-1b15-45ae-89aa-2c1521209817 · outbound

This paper cites an unresolved cited work.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.758759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.758759Z digest=sha256:114e850e916c36483a1adf4666ad79d8a730c4e3a5db696cc030a0452df88287

Observation c8f54277-54f8-4a76-8275-0f4db1c07298 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.769436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.769436Z digest=sha256:52495f92e3b48df741ec0056b536b4ea28194d582f29359983ae5c6dd7db3193

Observation d81f0077-0e19-4813-a18b-2556d0eb8c78 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.774679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.774679Z digest=sha256:c2ff5e77908b24a04d5c47b5903ba31e54914d16b2e90b5d577d32e022c2d55d

Observation a19b12ba-10c9-4aeb-b341-5ae3d70213c4 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.780850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.780850Z digest=sha256:f859295fe04bfc8b2c8cae4227f07310c23a8573bbc2e228b5f61c50bd1c7e6c

Observation d1fbcd2d-5eeb-4534-b26f-702e31d809bf · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.788297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.788297Z digest=sha256:48ce394f1a30ca07b7451c2f1a5020c9070f3d6c9b099e232fce85ba234530cb

Observation 3feb0f45-64a9-4928-aba1-01b0af72d850 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Generation and comprehension of unambiguous object descriptions

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.754295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.797421Z digest=sha256:06f77c6762020db4da5873285a8b7145186caf914fa8ddd5cad7cfa47fd02f3b

Observation c183780f-367b-400a-8d60-eec880502efa · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.734502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.803942Z digest=sha256:99ab028fcfac0099a3ab9b507975500778ccd8ed65d49f78ca21d714ab073ec4

Observation 205240e1-db18-4b72-8830-f1ed597a8fe8 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.815486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.815486Z digest=sha256:bcc54f146211a9e2d8e9e51f7470e716796dbed02476b2db5e38782065c13522

Observation c73cdfa1-7b31-4b9f-9cdc-5b6cdeee9f75 · outbound

This paper cites an unresolved cited work.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:47:01.715883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.820911Z digest=sha256:d59188e8ccb18e3c6c3fa6e4b11b5102fb44b7d75bc33ba9c4234ae08edc5a29

Observation 23cfa073-355a-4aaf-bacc-9af7ab0b407c · outbound

This paper cites Scene text recognition using higher order language priors.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Scene text recognition using higher order language priors

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.696298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.826429Z digest=sha256:e01f612c29913afd63ef193b9b40789003ba3aae18dfcdbe17611048a285846e

Observation 30e4065f-cd26-455c-8c27-ea349fe54298 · outbound

This paper cites Mishra, K.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Mishra, K

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.833230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.833230Z digest=sha256:302c51baa2f73f1cd6c5a148f39a5501436cf921137c6a64d7a3e21a9f6ececd

Observation dfff296f-3796-4572-b448-a7499d27f521 · outbound

This paper cites OCR-VQA: Visual question answering by reading text in images.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer OCR-VQA: Visual question answering by reading text in images

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.838784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.838784Z digest=sha256:0fe581dd8bde4f65a0a3792978136668de9428c52f3020798d5a98d38ba598d0

Observation af7366e7-23f5-42b2-9096-b4a98fa4865c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer DINOv2: Learning Robust Visual Features without Supervision

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.844180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.844180Z digest=sha256:cec9a7d3bf5239ca4d621517ade321adcfcfc0fdac1c7a7fc9372c233135ff13

Observation 6038b06d-71e0-4c7a-9447-ec297d7674db · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.649784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.849557Z digest=sha256:43602ff423e1e2032179e29a908ab0069db4d9121a47a2b21882172b93a13425

Observation bbad4ce0-ba70-4b7f-ba12-df7744ec0fc0 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Learning transferable visual models from natural language supervision

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.854568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.854568Z digest=sha256:0e2f6d44adb298258263f6bdbd7833e56e9405262864fb3f02bb25884edafd7e

Observation 87ce0e74-8d76-4f8b-8e4a-29170224c869 · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.860098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.860098Z digest=sha256:1c3b93b04e855a2f0dca2f0942cac51944e92805d91d86abf54fcabdb1d4364d

Observation 7e5157a1-3b5b-4027-8f4e-6df2cf4d1be6 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer U-net: Convolutional networks for biomedical image segmentation

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.619346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.866480Z digest=sha256:0f50ffd143e48eaed6f71bc043214e28f7b57ddc67cec3471af9f003393259e5

Observation 1ffbbb74-828d-4ef2-967e-98bc56d8650d · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer A-okvqa: A benchmark for visual question answering using world knowledge

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.596611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.872148Z digest=sha256:fcc3dfad7a372e38b0d736a5274f18c7d63b059cc110173efcda7661a9e75919

Observation c92c68b8-f2a8-488e-b768-ff2c6b798bd8 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.880614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.880614Z digest=sha256:e9e3970ae3673d1772d532ff8d86c4b21f188fd7d913fa688e1911b1753e89d1

Observation bdecb3c0-182d-4649-bdcf-6d1ebd23bd73 · outbound

This paper cites https://sharegpt.com/, 2023.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer https://sharegpt.com/, 2023

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.577922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.886130Z digest=sha256:be03dd23a6f80aab1d091f0a032523ec94e783d1812ddf975802fd23f116570a

Observation 6ab493a7-7d1a-4b78-ac95-00766ae2826f · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.893904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.893904Z digest=sha256:87a2159471610235d7d247c37ac2f4baddb6c5cbd18bdd86d4e1db2aea9d35af

Observation d416a56a-fc12-4a0c-a471-0d5edb227245 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Textcaps: a dataset for image captioning with reading comprehension

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.902180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.902180Z digest=sha256:8edd66df6fc08d5f69573a8e1e8ee2b17224ca08013fac1de5d1ebb991874e77

Observation 4749b436-e51c-4900-9937-13a9854f3f11 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.907264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.907264Z digest=sha256:05e3b506d8afe2c7962531b121f2e637680dea8443791a5dbfe5a0ff29fbf780

Observation 77deabb1-6b4e-4b69-ba1f-1533f34d298a · outbound

This paper cites Towards VQA models that can read.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Towards VQA models that can read

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.547774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.913563Z digest=sha256:6be5de1b8f8c67f7818dac077cd73ad43dab69a2304b73633c9396b96a8c4ee3

Observation 8026eddf-f93d-42e7-8f6c-3b3eaa0da6f7 · outbound

This paper cites Document collection visual question answering.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Document collection visual question answering

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:47:01.530032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:46:59.919338Z digest=sha256:3e90aa9b1456ed959775fa090bd11368555c980b35975eff3b0104e7419177ca

Observation 81ff6200-b0b5-4bd7-bfe0-fb453f77266f · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.925267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.925267Z digest=sha256:8c5122bd6abcf8e752287fcded668b588c901599b758adfb2a175e8b594b8991

Observation 5f3aa4d2-32a4-4ba4-8d90-210d7364700f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer LLaMA: Open and Efficient Foundation Language Models

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.931076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.931076Z digest=sha256:d5d225458346cb1344f64420073d5f15714edf0c35d4a0f4aec35e4957fd22d9

Pith citing papers

Observation 73a88199-3b83-4783-b056-416ef5d6628f · inbound

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding cites this paper.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.380728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.380728Z digest=sha256:0109a10cce54875cf97bd003c5c006b26445c07545ded6a41daa5b1b21a9744f

Observation 534df0bd-a894-4c44-9e34-44b76cf1d0db · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.321276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:71f2e8ee7e2f4aec1c623df9a310d0fc7004ec5c4e071c57f709a6316af55f4b

Observation 7bd86063-dc8f-43ee-b910-3d8236c3ae10 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:12.643287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:12.643287Z digest=sha256:634b21c264b0f8caaa0e8a456202f97695d6eb8a71d157d7126c416a9076edea

Observation 3d866024-da5b-4073-b695-8bf8b0b3a9ae · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.447271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.447271Z digest=sha256:66b74c36977a80b30def3a8e564d2c0badba6362d24f28616084b802738a62bf

Observation 90a9cec5-14ed-48b0-9266-55363b3a1af6 · inbound

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints cites this paper.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.575266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.575266Z digest=sha256:070bdd5a3f998dbf719a125344fc1bd73ed58a756c2cfe51d702b2b127b643d0

Observation 3f7e994b-b245-48c3-bdd9-998ef8bb5e91 · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.766319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.766319Z digest=sha256:5e53d3718aa3de05b36505d9d4d2393f6d5b1f1736098dbfac72ae26b891cf25

Observation 43792357-ef21-498d-85c7-8bdd247bec22 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.975128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:976bd263c625e4c680fd3f034c1237bd111037e8de0d4a4dab654d6fc14b3e95

Observation 326c4e2a-52cf-4612-b9f8-7edb3c605d47 · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:52.954889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:08209b082b93eda74c1f60163670da7f97753cae3fd9e7bf9547c95934e93502

Observation cdb03c63-ac96-41cd-b557-54e2890f3a6c · inbound

RADIO1D: Elastic Representations for Condensed Vision Modeling cites this paper.

RADIO1D: Elastic Representations for Condensed Vision Modeling LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:a57cd037933ffd7ea4844b6a36ff11e082eeba204573c0ad20649d8585341a58