Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:18:49.356798Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2501.10318.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:18:49.356798Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T18:03:53.318565Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-09T18:03:53.492309Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4893557c-b3cf-4b9c-ba44-78d3b3d48105 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a7fa20c-a563-4222-b8b0-0bff5bec1197 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ba65c0-b72c-4a44-ae2b-4c89cbf7cb93 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151e497f-0dc5-44b6-88e5-e4d8d25dda27 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e44bf7a-98fb-44d8-85db-97d940774865 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c625dd2c-4282-4b21-88f6-1bdc088c653c · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Introducing our multimodal models, 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8c1d574-848a-46c5-a34f-a75bcf6a1649 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Matryoshka Multimodal Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6168ff7-5972-4cd9-a44e-b15707bbdf69 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0ba5e46-5fff-4451-8d9c-21aac81f372e · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1bd42b4f-011a-49ca-b2d3-365fceb83c3a · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 84081bb7-d1bd-4704-91ac-1f31a5d53b5a · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 580a9ceb-4d1a-46af-92a1-af0e7d9536d1 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce90f7d8-daa3-4337-84c4-5846203e3238 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c8ea428-d2c8-4e8e-81bd-b020fff53a28 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c52b67d-40ed-4e67-a665-13d683f6c0ca · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4523bc-06f8-4b7f-b16f-39d5ad5a2678 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e552329-6299-4b6d-811b-5025bad163b5 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 58f2dfe0-49c7-4246-9b9e-b23a70f6a75a · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7b034409-e9db-4e28-9444-bbfdec447f22 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f38d5e1-2401-44d9-84b7-7033e945acb1 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb438ae-dc3d-4ee8-a6e4-7b56ba76931b · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb9e5f6-930c-48f3-9795-fc1b81ad1dd9 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 533e2edb-6634-435f-a09b-2c7f9bc28248 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3eb2d70-b7de-4446-bbe6-dacc3b43adf4 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Improved baselines with visual instruction tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5a72f689-464c-4052-8ac2-c6e3fff0326f · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Visual instruction tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72a9111b-b727-49c9-bfa7-510299d39dab · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c33ae2a-4435-4a3e-94e8-f0d685e0597c · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Towards vqa models that can read
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5cb256b6-a32a-4743-9e5f-0f4d058ad732 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Gemma: Open Models Based on Gemini Research and Technology
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf04df61-7787-48d4-98ed-2e439ed84300 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c93814c1-7c05-4f13-a946-d2dc4e07d377 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Tinyclip: Clip dis- tillation via affinity mimicking and weight inheritance
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b06da072-b8d1-4d28-b129-5ebd99ff7f3c · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models calflops: a flops and params calculate tool for neural networks in pytorch framework, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 117951dc-7cd1-49e9-9289-b655767bf5a6 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Qwen2 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfb0970e-b584-44be-8b15-b4885abeafbf · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbbe59b8-c29e-4934-893c-205a33ddf6c7 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 15ec2280-2c26-46c3-96ab-623a1531fe47 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models GLM-130B: An Open Bilingual Pre-trained Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05e6d4d-d2d8-44e7-b630-6a91045b1c73 · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models Sigmoid loss for language image pre-training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f430555-2d21-4757-b227-c0aedde912da · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models TinyLlama: An Open-Source Small Language Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8081ebb7-4195-4af3-8b03-f4a8d42845ee · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 703216a1-543f-4bc5-9901-d188a215f31b · outbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1837d5-e9a4-484d-a90b-0c32bec08abb · inbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework HiMix: Reducing Computational Complexity in Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e8af6555-174f-49e8-9fb3-4932eeb61e6d · inbound
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models HiMix: Reducing Computational Complexity in Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.