Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:25:00.399099Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2411.17773.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:25:00.399099Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T22:40:39.892802Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T22:40:43.026280Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5230cde1-3152-4f7c-b770-1f9f6c3b7551 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef091b08-659e-444b-a3ad-8becc7b819e9 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2e6de39-00dc-4dfb-8d64-9cea8fe19938 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Improving image generation with better captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afd818f9-a137-4b25-88c2-cf1002a6f4b2 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Token Merging: Your ViT But Faster
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64644470-221f-407f-89bd-3744c2d2c8ee · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Honeybee: Locality-enhanced projector for multimodal llm
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd12837a-9450-408b-9fd3-ad712256895b · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d47b8885-67fb-41c9-b767-57ae6ed2ef53 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b1a71e4-03c2-48aa-9ae9-0176bd548584 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09893b2c-e01d-47ac-9d3c-3cc20284bdd4 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 214b2d25-4901-43fd-bf37-467ef56475b7 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a977841-9611-4c50-928e-2d65fd479c4a · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b90320-8df4-49bd-b9f6-45ecfd133974 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60001f2-cd20-4437-804b-4f2d2c99ec73 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4b2d2c-4d3e-4d45-874d-7c293341aebc · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Evaluating Object Hallucination in Large Vision-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0358a844-7806-4255-9a7c-cc027c1d2bf3 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0442edb4-31e6-4af3-84aa-3c588239285b · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c04a221-ead8-4f7c-9d06-9c06c01835e9 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Visual instruction tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c3ade4e1-85ea-4eab-89ca-8399625fe3f9 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping MMBench: Is Your Multi-modal Model an All-around Player?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c8ee26-7b4e-4e2b-a9b3-98e9b32fed22 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405752d7-7327-4b50-a1a7-d897cf44dcd3 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Learning transferable visual models from natural language supervi- sion
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a07fea6-86ee-449d-ae0e-ece8976773c9 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Towards vqa models that can read
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13052108-b106-43fe-8134-5a5b9cf59b3f · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6956228f-437a-49b9-924a-2f03b47292b4 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d9fce147-c50d-4345-a667-700ba991a541 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd63667-5baa-4b2c-a5ee-0bf1ad13ba94 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Neural discrete representation learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5357ac85-5fa8-4cb4-9c01-a3e34be52516 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b66a686-60c9-43f5-8695-8dfcd767f616 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Efficient Streaming Language Models with Attention Sinks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df0f32f-6fbc-4d4e-b8ea-51feb0162fd8 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Groupvit: Semantic segmentation emerges from text supervision
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5ecc5ec3-ec4b-4125-a631-81b5b5c457dd · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0cb8b0-25e8-4e56-a9ff-a24c72415139 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49af9acd-90f2-4ab2-8674-a7dcf84265aa · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping VoCo-LLaMA: Towards Vision Compression with Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e842cc9-a5c7-4467-b472-628f30265967 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a9a7bdd-8385-4453-a25e-79d65249bb9c · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping Sigmoid loss for language image pre-training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6806a50-cea7-4973-82db-3383cef06e19 · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0471f26e-8674-48b7-bf4f-07ea6732e36c · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c17ded9e-09bf-4628-a2fd-3688b23fad5b · outbound
Efficient Multi-modal Large Language Models via Visual Token Grouping MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87aad0bf-2f97-415f-9013-4e57a68f7c86 · inbound
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models Efficient Multi-modal Large Language Models via Visual Token Grouping
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.