Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:46:59.931076Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 9 inbound Pith citation observations for arXiv:2412.13871.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:46:59.931076Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:39:35.380728Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T09:43:00.317941Z
100 of 119 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5d938658-7d66-4962-baa3-f104b23cdbad · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer https://huggingface.co/datasets/xai-org/RealworldQA
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8726ff5f-9218-4400-89a3-342f28550a86 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer https://openai.com/blog/chatgpt
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89178c1f-e3ba-487b-a310-eaf481564107 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer https://huggingface.co/datasets/laion/gpt4v-dataset, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cebf44c4-247f-4149-a860-7f10a008f760 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20896bd-eb2c-4f37-b4a3-c15722d4b43e · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Understanding intermediate layers using linear classifier probes
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9bdd3b-9aa3-488c-91f3-2bd087fd22d3 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26f5bef-9fa9-437d-baad-921084d76ef2 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Multi-label cluster discrimi- nation for visual representation learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328a96cf-04c3-4bc5-9b45-c36c00a4d7d5 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Gemini: A Family of Highly Capable Multimodal Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12af006f-8cf2-4386-84e3-c77c47027708 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer VQA: Visual question answering
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 440a624d-a99e-476b-838c-c8cbc026f9c8 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Vision transformer for fast and efficient scene text recognition, 2021
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a81f5960-63ae-4a89-808e-bfbd42309127 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03e2fd5-77dd-401e-82ef-045ff8e0ce75 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Introducing our multimodal models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2840d8cd-8db7-4540-9e21-53c2ad50dee7 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Token Merging: Your ViT But Faster
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9beabfef-632d-4ce8-ab61-ace1893a7885 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Burt and Edward H
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1bfaab-7b61-43a5-b432-6ed8d8abb769 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18ae7ab-6da3-4f44-b5c9-9058b371e453 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Honeybee: Locality-enhanced projector for multimodal llm
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ecb40d-8b52-40eb-8f63-cca626689cc3 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2efd09-6a65-4065-bb2f-bb74d5075251 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af2d6da7-065f-4bad-873b-d9af5e5e3b80 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40f7086e-4021-4ef2-9a6d-35cf011e9801 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6400ae54-8e11-4b0e-974a-666d9771c0b1 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d27094e-ac80-4478-9cc7-737b3866e3d9 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73e02ea4-56cd-4aeb-be70-1336a6b14ac4 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81438696-186d-4102-b4a0-5c48293d3d70 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Xtuner: A toolkit for efficiently fine-tuning llm
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b02e12-757b-45f5-823f-fd29447ac61a · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb885f6b-a024-449c-a5b2-f291cf5bd7e3 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e57181-18d6-410e-af4b-e35b9ffeb216 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer An image is worth 16x16 words: Transformers for image recognition at scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8e2d69-78ca-461b-93c9-914821742653 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c72ca9-ce90-4dfd-b677-93297f6c9d68 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer FeatUp: A Model-Agnostic Framework for Features at Any Resolution
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e369f5a0-d97b-49f2-b739-e8b58a348070 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f912bcb-b483-4bd9-8eef-d73b36f1e3da · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6000433c-cf29-4f0a-9898-d7fc1ee58c29 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d242c4f6-fc85-4337-ab09-75c0a09799d2 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bab988e-6a6a-4d8c-88ca-92e65352d052 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a978c192-42c0-42ee-9b53-13584642a8fc · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer LLaV A-UHD: an lmm perceiving any aspect ratio and high-resolution images
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464e57b3-2e21-4c87-a474-64c4a3faba6f · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer VizWiz grand challenge: Answering visual questions from blind people
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a326f9-93c0-4c86-a0d9-db68247c9504 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer It Is Likely That Your Loss Should be a Likelihood
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57fa13cd-02f6-4dcb-b617-ccf4beada08f · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4534fd55-d7ee-4fe3-ba40-675990ad0658 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Deep residual learning for image recognition
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d745fad4-e8ee-4ec4-bc30-0a3c9bfed479 · outbound
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c829448-aa43-4727-aa99-45ef772a8d61 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer CogAgent: A Visual Language Model for GUI Agents
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b99463a-c2f5-481b-9225-59fba3f95ded · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer GQA: A new dataset for real-world visual reasoning and compositional question answering
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9610f394-12d8-4f88-8063-18e84da21ee6 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Synthetic data and artificial neural networks for natural scene text recognition
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53fefc95-47de-4e49-ad12-57cc8dfd8bbe · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ab31ee-d5f2-4eb2-ac77-ed06829912ac · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer DVQA: Understanding data visualizations via question answering
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e61f89e6-39e6-49d9-ab70-abf183630820 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Referitgame: Referring to objects in photographs of natural scenes
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf0c10f1-b32e-4690-9f68-1db2db626a10 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer A Diagram Is Worth A Dozen Images
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a84b43-14ec-446c-8aa3-1ed65a65ff0b · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer A diagram is worth a dozen images
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844a8005-5df7-4d1f-a926-ccf5b365c54e · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d4c8b8b-f4a1-41b8-9c54-51e05732990e · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Ocr-free document understanding transformer
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccebcbdb-3519-4b9c-8dae-d85385937726 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Segment Anything
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3300d09-90af-478b-a43a-7b6499cde8d3 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7adf03-c513-4f38-89c0-8eaf96a2b13f · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Pix2struct: Screenshot parsing as pretraining for visual language understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebdcab75-016e-4367-8fde-6652120f67dd · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95274e59-dab2-45e1-8daa-0e440bc6809e · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer OtterHD: A High-Resolution Multi-modality Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2570af2f-4935-46e7-bcce-0fcfa510b934 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb05bf86-a2c6-460a-8300-fb344eaa035b · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c1b57d-9700-4cc1-bf89-c0f4903b85ab · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f99a26e0-6a84-471e-9778-0f011d2d64ae · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fabb162a-e7d9-4bd3-a0f6-dc13ef9390b1 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Vila: On pre-training for visual language models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a0cb0c4-3b8b-4c2e-ae79-f1ebbd2aa2f7 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Microsoft coco: Common objects in context
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f17a4449-0be6-4843-b37d-470c68708fc9 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Feature pyramid networks for object detection
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a28a15d7-6341-41e9-9f43-9fd5a5a1b2c0 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a07460a7-c114-436b-b1ff-0d8207bc7a00 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a415545-e2e1-4593-8ba9-e0cca7385aa9 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer LLaV A-NeXT: Improved reasoning, ocr, and world knowledge
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6964874-820b-4799-b1a9-182c9f0838a8 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Improved Baselines with Visual Instruction Tuning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130ca913-574d-468a-9e4b-3dcc5753d830 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Visual instruction tuning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adef2197-0dab-46d3-a36c-e5e5a1026e85 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MMBench: Is Your Multi-modal Model an All-around Player?
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700232e4-a985-43e0-9b14-21cae2eeb0d8 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c8882b-a001-4348-bb7c-b9d814c12715 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 570199d5-1f0c-4c73-b2e3-9974e24a9c9b · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Swin transformer: Hierarchical vision transformer using shifted windows
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7927ca1-ff8f-4ae2-96ca-87fe29363896 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer A convnet for the 2020s
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d2e19d-f8bf-4f2d-b833-61b343c53b28 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e9c1e5-1b15-45ae-89aa-2c1521209817 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Unresolved cited work
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f54277-54f8-4a76-8275-0f4db1c07298 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81f0077-0e19-4813-a18b-2556d0eb8c78 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19b12ba-10c9-4aeb-b341-5ae3d70213c4 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1fbcd2d-5eeb-4534-b26f-702e31d809bf · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3feb0f45-64a9-4928-aba1-01b0af72d850 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Generation and comprehension of unambiguous object descriptions
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c183780f-367b-400a-8d60-eec880502efa · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 205240e1-db18-4b72-8830-f1ed597a8fe8 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c73cdfa1-7b31-4b9f-9cdc-5b6cdeee9f75 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 23cfa073-355a-4aaf-bacc-9af7ab0b407c · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Scene text recognition using higher order language priors
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 30e4065f-cd26-455c-8c27-ea349fe54298 · outbound
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfff296f-3796-4572-b448-a7499d27f521 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer OCR-VQA: Visual question answering by reading text in images
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af7366e7-23f5-42b2-9096-b4a98fa4865c · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer DINOv2: Learning Robust Visual Features without Supervision
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6038b06d-71e0-4c7a-9447-ec297d7674db · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bbad4ce0-ba70-4b7f-ba12-df7744ec0fc0 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Learning transferable visual models from natural language supervision
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ce0e74-8d76-4f8b-8e4a-29170224c869 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5157a1-3b5b-4027-8f4e-6df2cf4d1be6 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer U-net: Convolutional networks for biomedical image segmentation
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1ffbbb74-828d-4ef2-967e-98bc56d8650d · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer A-okvqa: A benchmark for visual question answering using world knowledge
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c92c68b8-f2a8-488e-b768-ff2c6b798bd8 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdecb3c0-182d-4649-bdcf-6d1ebd23bd73 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer https://sharegpt.com/, 2023
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6ab493a7-7d1a-4b78-ac95-00766ae2826f · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d416a56a-fc12-4a0c-a471-0d5edb227245 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Textcaps: a dataset for image captioning with reading comprehension
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4749b436-e51c-4900-9937-13a9854f3f11 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Very Deep Convolutional Networks for Large-Scale Image Recognition
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77deabb1-6b4e-4b69-ba1f-1533f34d298a · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Towards VQA models that can read
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8026eddf-f93d-42e7-8f6c-3b3eaa0da6f7 · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Document collection visual question answering
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 81ff6200-b0b5-4bd7-bfe0-fb453f77266f · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3aa4d2-32a4-4ba4-8d90-210d7364700f · outbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer LLaMA: Open and Efficient Foundation Language Models
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a88199-3b83-4783-b056-416ef5d6628f · inbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 534df0bd-a894-4c44-9e34-44b76cf1d0db · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7bd86063-dc8f-43ee-b910-3d8236c3ae10 · inbound
Reinforcing Video Reasoning with Focused Thinking LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d866024-da5b-4073-b695-8bf8b0b3a9ae · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a9cec5-14ed-48b0-9266-55363b3a1af6 · inbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7e994b-b245-48c3-bdd9-998ef8bb5e91 · inbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43792357-ef21-498d-85c7-8bdd247bec22 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 326c4e2a-52cf-4612-b9f8-7edb3c605d47 · inbound
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cdb03c63-ac96-41cd-b557-54e2890f3a6c · inbound
RADIO1D: Elastic Representations for Condensed Vision Modeling LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.