Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:21:34.883733Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2411.15453.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:21:34.883733Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:19:24.219539Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T17:45:25.594440Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bfdcc233-3c64-440b-a9ba-6e6676a59393 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Transformers are Multi-State RNNs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec0190d-6334-4ec8-ad40-09fbc9f4dedb · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy , " * write output.state after.block = add.period write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97481105-b6bb-4edd-9482-468496acff33 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy write newline
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb39386-d352-4df8-b5d3-25d4ec998ebd · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2883594-10bc-48cf-a758-52a2f3ecf323 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8ba58933-ebfb-4b45-b033-fc65a2a9d328 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Qwen Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 934d00c2-332b-4d18-b965-c404f8669cfa · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2b603d-3e20-47e0-91c6-e83a0827ee7f · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd31cc0c-f7a0-4a03-93cd-82d5f2a7d105 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy PuMer: Pruning and Merging Tokens for Efficient Vision Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033d3e43-1c10-4d41-b9d9-17738af2932d · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 220e32d6-2416-4f56-b817-71c22e3d438d · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a91a72e-b111-40a8-81fa-d69822036f7b · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5677ec-ce38-4ff4-920b-d4aa6a29049e · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy E.; et al
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476d3616-7fa3-47fa-8be9-066cf295c02c · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d81df49-44a7-46fc-be70-918bee1e5dc5 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 33a1fe99-d3d3-4df9-bc2e-4b8e59360e68 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa32e8c2-2e75-4a29-98e7-558b34d29205 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8d4602-3f82-44d5-a6cb-1e098fd4b37e · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5994cfb4-0c2b-47fb-b066-1e84d08e9b1f · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8fcd0e5a-a5c4-41c0-8003-fef3fe760a52 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6487c2b4-59f1-4538-aa4e-0767c51e811e · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Gaussian Error Linear Units (GELUs)
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 794c8a28-881d-434b-afe8-b640b4a62ed2 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy A.; and Manning, C
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28dd39a-3662-41fa-a368-26501a2b1cd1 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c0984625-26e8-483a-8e0b-cb1c6682a783 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy AI Alignment: A Comprehensive Survey
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cde3319-c7e3-47de-b200-b2d9fb282b7f · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy M.; Bommarito, M
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 304945a7-11f1-43ba-8ea5-4def64e1f41d · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da516c66-8aba-4538-bf4a-758f49d5f36d · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d79ab744-ba17-45b0-9e0b-ab8799d6b314 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy OtterHD: A High-Resolution Multi-modality Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f9b2a6-da60-4150-a6a5-67274334256c · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d112de80-b060-45f1-9afd-c2a66e8717f6 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fdc6136-e95c-4799-b64c-b594606e0d99 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy VideoChat: Chat-Centric Video Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d52d78-545d-4e99-b1cf-c99881318b9b · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e3ce87b-2785-4118-a9e7-ee615e3344d7 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aade3e33-5607-442c-96a5-b7dcc3742a0a · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 829a9214-bafc-46d1-b88c-9488f25397a1 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc045fd-3104-44cc-97be-ab789dbf8c3d · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy MMBench: Is Your Multi-modal Model an All-around Player?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae51f99-e8c4-43d4-9e4e-e4dd83cc4877 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b880f91c-ec04-4505-b538-480c237f60b2 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40649477-08fa-4b19-91f8-e22bf1433a3e · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 505bd8f3-378d-4145-9875-c8cc1c52adbf · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3d6a5d-b454-40fe-a1c3-d50bb490a486 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5942c393-4a04-4fda-9c5a-cad874f379ba · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f55ba0f-5e3d-4ac2-9a7b-5e88cbe8d15c · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy J.; and Yan, Y
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acbabdd-d441-4abf-a6de-c4091218ef70 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19be0be-6dea-4578-953b-8f88991e100f · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a64b03-b089-42a2-9944-5a22f60b2b84 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy LLaMA: Open and Efficient Foundation Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634f3212-6bf5-41ba-b408-0c6bca286042 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c9044f-8546-4431-9d3b-f534c47f891d · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Attention Is All You Need
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0e6642-c249-4dca-8626-087a0eddbcd1 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe18d07-90c8-4940-9f68-66e75506af6a · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b7c15e2-fd0a-4fac-99ba-3837589499df · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Should You Mask 15% in Masked Language Modeling?
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef32f889-edb0-484f-86f7-91a1467a4adb · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51149963-7a1f-4e81-a36f-2f4b57600d24 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d71aea1f-e280-4820-8845-7debb73c14fd · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84837abe-8bbc-4d0b-9e66-bfaa44e232f9 · outbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Instruction-Following Evaluation for Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba3da3bd-99a7-4c67-87fa-f21e8ebb26f3 · inbound
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f681dd-5629-44b2-99c5-05cccb89b915 · inbound
GoVector: An I/O-Efficient Caching Strategy for High-Dimensional Vector Nearest Neighbor Search Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.