Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:16.136813Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2505.11945.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:16.136813Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T07:14:08.479610Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:08:21.980388Z
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f62878cd-7277-4c8d-a45e-65f364e40591 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2f9599b4-91f2-4f73-a3f4-4b4f9829b513 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0861b875-b5a0-4ac6-8897-3dfa2c7dd57c · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1e6f6c-b87d-48a1-8d1b-2ff46a275d0c · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Honeybee: Locality- enhanced projector for multimodal llm
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5e5dd523-61d0-4079-9fc0-db871464aa34 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f9dee378-b90f-4403-a288-3f01828d9a88 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d2fa2d2-90f0-4c95-bf16-5145408b2b86 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d3af01d-f0fb-4eab-bff0-c6108479f7a8 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f1ecc6-94c2-4125-a27b-15255b2174bf · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Don’t look twice: Faster video transformers with run-length tokenization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d8f9a9cd-0d68-487f-b0a9-7502e8ef8f87 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b2899d-2898-49e0-aba3-f4e768f28024 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0df83c2a-7098-47dc-8e16-ab180eeb1650 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e9f6d2-3058-4da3-91a6-df0fd278f016 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 16cdcd33-faf0-496f-96e7-ee146fc3d012 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40832420-5bc7-414a-96e0-626514df2af2 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Efficiently modeling long sequences with structured state spaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 98a84017-9be8-4346-9658-a016bab80d0a · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 58d24812-16da-4f2f-8774-ff66ed7f39a4 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vizwiz grand challenge: Answering visual questions from blind people
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0cf33c95-b05d-45aa-a327-2103b752530d · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Bliva: A simple multimodal llm for better handling of text-rich visual questions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 055fb796-142a-4db1-9318-de1e4de9a968 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d026d179-89cc-4830-9ce7-8dd29369ae1e · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Token compensator: Altering inference cost of vision transformer without re-tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cee4edba-c37a-4e81-92ca-93dbacb28d72 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Logicad: Explainable anomaly detection via vlm-based text feature extraction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 56fef084-41e6-4492-88a2-c0fbfcbcb10a · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c870c746-a0ec-401c-8f2e-ef8df16d13c7 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40462433-8ad0-4bf1-a6ce-6d052dd65f80 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48131032-88e0-498d-b4d1-dba3fbf890d3 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Lookupvit: Compressing visual information to a limited number of tokens
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 16f1ee93-5146-4761-b480-e14b0c4e315a · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c32cae-143e-4ffc-9aee-35d91ae85dbb · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3601f0f1-856f-4326-8837-e95d082e1208 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d080d537-2f39-4dff-b5bb-e43875d8d4e6 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832cf5ed-d8b9-492e-921a-8ff379890722 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f1754f53-f7dd-44a1-b20f-3ddb0d75739e · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 82cf2356-dfab-4ab6-ac15-a58e88dd345e · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b7a0790-66f6-455b-a4f4-4dedf1e16902 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb0d0afe-d458-4f1f-8384-fc77cc1ea76f · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llama-vid: An image is worth 2 tokens in large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 204187f5-f6e1-456a-8ead-79c436535559 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10ef480-4df8-47ca-afc9-4380f5a04b49 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Evaluating object hallucination in large vision-language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f4aab4f0-8be9-4c5d-aee7-460d243a8e03 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Monkey: Image resolution and text label are important things for large multi-modal models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c32f603e-a5f6-491e-9037-dc2f2dfe2e2a · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vila: On pre-training for visual language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 12e92281-0c0f-4df5-b138-a885811ccef5 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e34a6ef-d5ad-49b4-9c7e-9fbcb4e6a523 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Improved baselines with visual instruction tuning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b195c883-cf1e-4be7-8117-1de4ed7fd9f6 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llavanext: Improved reasoning, ocr, and world knowledge, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b99e44de-561b-47f8-bb64-419825bc6dc4 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual instruction tuning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2af17887-9523-4eef-99ba-52009acf4f1a · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b9be14-69f3-4160-b451-211b1d53e57a · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f529d4dc-f324-43c3-9a9b-4adefaa2e53f · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d473a4a-7a10-416e-9a8a-d5879019c480 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Questioning, answering, and captioning for zero-shot detailed image caption
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ed2d803-7c77-41b0-a3a6-b31c97e84256 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6b057137-b19c-434d-9e52-c74a4c018f8b · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33740278-dc71-4c6b-a779-fae282c58dc3 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Infographicvqa
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 44c17a7c-f79c-437d-89f4-eb8960cce2f0 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Docvqa: A dataset for vqa on document images
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 404e1c35-cf10-4e80-b998-2b129e6b152d · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ocr-vqa: Visual question answering by reading text in images
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a17bc5c5-59a7-4933-b169-1d602e105c20 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c172126-ab99-41b1-96c4-0fb76ffaa218 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Learning transferable visual models from natural language supervision
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524c67da-e1ee-419d-8dca-9bef68068679 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f748f79-0d26-435b-8d19-8eed9c027152 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Textcaps: a dataset for image captioning with reading comprehension
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c46163-dada-4f41-a7ff-6e278ad53d2a · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Towards vqa models that can read
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc2e32c-88a0-4a0e-9659-ace1bc2b6246 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Less is more: A simple yet effective token reduction method for efficient multi-modal llms
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 77bb4be7-8de6-4b4c-9f86-11a44302791e · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c41836-a8eb-4564-bef9-77486e1a450e · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FastVLM: Efficient Vision Encoding for Vision Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d4f5561-1963-4097-8a5c-13ca50a36610 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2a1795cc-d234-44ba-9259-e91f96f51404 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Fashionvqa: A domain-specific visual question answering system
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ae314ba2-7d11-4713-9bc4-f20418aa3413 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Cogvlm: Visual expert for pretrained language models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9ba78d9-d7b9-4fdc-9969-9a291cf2c333 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Rl-vlm-f: reinforcement learning from vision language foundation model feedback
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8610864-5872-44be-8ce8-29b904cb20d5 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vary: Scaling up the vision vocabulary for large vision- language model
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31167906-34d2-4242-b1e7-f9d258976af0 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19fe73d8-5170-463b-81d9-29179927abc9 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visionzip: Longer is better but not necessary in vision language models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dce44ee-3c80-4486-8fa8-017cc21a1f7b · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39af3fc9-e525-42ce-b7e7-76b20d873fbd · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c5d5a9a8-76c3-467a-ba48-ed8a6b1fbd4f · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5b262117-94e1-4ebc-a208-b54d27729888 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e4b296a-43fb-4a93-82f4-ad1d4a8dabd1 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-mini: Efficient image and video large multimodal models with one vision token
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec64c60e-2bcd-4618-af58-83360ab3a170 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6ac840-ab06-4b33-b0cd-7d4836c0c254 · outbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbac5bfd-94e9-46e3-a2ec-9a76cecb0132 · inbound
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0403b79d-6273-4590-8527-cf80380382d0 · inbound
The Hidden Power of Scaling Factor in LoRA Optimization Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.