Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:46:07.815623Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2506.07138.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:46:07.815623Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:18:00.869557Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-14T22:18:04.010679Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 51da7a16-600f-4cdb-aa5e-df1301accd5a · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Gpt-4 technical report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 242480fd-5e69-4532-937f-d1e28933de3d · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Mme: A comprehensive evaluation benchmark for multimodal large language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 47b7ccb5-ecec-4232-aa66-fdd0e4a612cf · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Llavolta: Efficient multi-modal models via stage-wise visual context compression
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 801ecf1d-d2b5-4ccb-8796-7d3024829c01 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a5725ba3-bb26-4992-8729-7bca59097c77 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2daf0eb1-0da0-4355-a70c-aa749f051295 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Funnel-transformer: Filtering out sequential redundancy for efficient language processing.NeurIPS, 33:4271–4282, 2020
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7bd3348e-dfab-4b74-9a06-0e3d3a1b3da2 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation daba7e17-70db-4bb9-bec1-7ae646e68133 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Eva: Exploring the limits of masked visual representation learning at scale
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d92b89d5-1e62-425e-b63f-2aa2053392e7 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Sparsegpt: Massive language models can be accurately pruned in one-shot
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aca5367f-5d25-4ca1-a0fe-0fd6e65d4311 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d5bca0a3-8e1d-4418-af6a-33bd85c37fc3 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8e06c92b-03c3-444c-a377-47214adbdf06 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Vizwiz grand challenge: Answering visual questions from blind people
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4769ae5f-cd76-453f-bbed-d41799705769 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bfaac42d-eeec-4e09-8532-e67ba4abd52b · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Gaussian Error Linear Units (GELUs)
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e66b3781-7f0e-48ba-ac60-57e40cc8b359 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 950b0f9c-fc84-421f-8560-abb520e05b58 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Phi-2: The surprising power of small language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e73ff6e0-9ac9-4781-8b96-d51dee059a30 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Tokenpacker: Efficient visual projector for multimodal llm
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e7917c06-20eb-4107-8422-23d6d2c4e03a · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Mini-gemini: Mining the potential of multi-modality vision language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b47089e5-752d-45e9-b401-9ac7e88e70e1 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Evaluating object hallucination in large vision-language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09630c67-f9b0-4bd9-bca4-aae99b6e66cf · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Improved baselines with visual instruction tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5b7854e-9639-456f-b092-1c9da409e187 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Visual instruction tuning.NeurIPS, 36:34892–34916, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3230107c-e0a4-43c6-a16d-116a70d0056b · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f77dbf56-8bcc-49c9-b43f-de9eea84a03a · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Learn to explain: Multimodal reasoning via thought chains for science question answering.NeurIPS, 35:2507–2521, 2022
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c35a9f06-369b-4659-9e0d-b75570a45183 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Are sixteen heads really better than one? 32, 2019
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a5937e6c-aeca-4516-acf9-e22563477cc9 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Efficient transformers with dynamic token pooling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1974f68b-547a-4176-ac2b-2e6379512855 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Learning transferable visual models from natural language supervision
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65373317-e8cd-4802-be78-97e2c620473b · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c900fbc-8a38-45fb-8e84-9bc63e91248f · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 34caa49b-be5c-4abe-ac0b-bb8a93e3d73a · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Inter-Instance Similarity Modeling for Contrastive Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9472e061-3934-41ec-b802-547a485a9f57 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Towards vqa models that can read
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd25046-4cd4-4547-bf4a-a9afc2d8984d · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Data-efficient multi-scale fusion vision transformer.Pattern Recognition, 161:111305, 2025
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e4993ce-e0cb-48f2-82c4-8bb6a410038d · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Gemini: a family of highly capable multimodal models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 39e2a535-6473-4740-b768-40ed8d323bdc · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Llama: Open and efficient foundation language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e5e1a986-84f4-4c31-9c98-24dd65ce52d2 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Attention is all you need.NeurIPS, 30, 2017
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aec081c-b095-4e6d-82fb-aba3f29c9209 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Calflops: A flops and params calculate tool for neural networks in pytorch framework, 2023
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 20d354e2-bcc9-4296-84ce-84320e6d704f · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Texthawk: Exploring efficient fine-grained perception of multimodal large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9f0b0680-df8c-4f28-beba-fce3e75b9e52 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Vcc: scaling transformers to 128k tokens or more by prioritizing important tokens.NeurIPS, 36:20260–20286, 2023
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e8199289-00c0-46d1-8199-3dfd759be749 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Tinychart: Efficient chart understanding with visual token merging and program-of-thoughts learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 45577e7b-f0ee-4765-8fb5-8d49e649f5e2 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Llava-mini: Efficient image and video large multimodal models with one vision token
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a8f7b169-0b02-4f72-a8ee-6909f489d59f · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a30af405-ec88-473f-a549-af8c0623ceb9 · outbound
Learning Compact Vision Tokens for Efficient Large Multimodal Models Treat visual tokens as text? but your mllm only needs fewer efforts to see
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c71ed3b0-259e-4ab7-9df1-52cbaf754290 · inbound
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models Learning Compact Vision Tokens for Efficient Large Multimodal Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb49eef2-ff18-401d-a49c-9f8578789311 · inbound
Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression Learning Compact Vision Tokens for Efficient Large Multimodal Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.