Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:00.817362Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2501.08443.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:00.817362Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation be73d4bd-ef57-42ee-874f-f9a60d38f61d · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Improved baselines with visual instruction tuning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8490bfac-f1e6-4a41-aa93-ec56dabbd854 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d701c504-7e14-4aa5-8d31-d043187974f5 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89922604-efeb-46ff-9162-fa71417b4beb · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8d25a3c1-cc19-4567-a9b9-fb99a4f20347 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Talk2bev: Language-enhanced bird’s- 29 eye view maps for autonomous driving
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf590dee-02c1-45e6-80ce-e2caa7310f03 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models What can VLMs do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition , 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2fef787f-f3b7-4cc9-b564-136d50c348cf · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a057600-aee0-43cd-aabc-e2ed182e82ea · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Learning transferable visual models from natural language supervision
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74fb2fdd-49a9-4c36-b160-f9baa2d2b086 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Sig- moid loss for language image pre-training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e70968f1-08fb-4f95-9d73-b04acc9a92c0 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models A Survey on Hallucination in Large Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c60e1f1-274a-4bfa-914c-0e559bffa61d · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 653fdad4-a124-4cb3-b982-d97938d745f8 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Teaching matters: Investigating the role of supervision in vision transformers
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af3d2fe9-75c3-4303-b35a-1ba8b32c0ed8 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models What do Vision Transformers Learn? A Visual Exploration
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f23e13-0e69-404d-b768-e6734529e72e · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Dense Connector for MLLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8246b6d9-140e-4581-9051-2b84db952a05 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4631b653-1938-4e6c-b03f-3a27ec22ea07 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 898b8b5a-0402-4615-9ee9-2154ab66256a · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MoVA: Adapting Mixture of Vision Experts to Multimodal Context
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3580f57-c49a-4d0c-a359-49cfc8d51145 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d7f11d-e134-46bf-b1f3-369914cb3555 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models V?: Guided visual search as a core mechanism in multimodal llms
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 416e3d2c-97ea-48a3-b3b1-1155cfbbbe3c · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02f5a5f-4c2b-4d89-96db-5a0f22b816c8 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b31f994-3db9-4df9-8637-ec7370a7d065 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdfdb95f-bec3-4c6a-ae64-c86047a500fd · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Gqa: A new dataset for real- world visual reasoning and compositional question answering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 33bbe69a-0068-40b3-8dc3-4d50962adb11 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Seed-bench: Benchmarking multimodal large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2ff97242-887e-4944-a538-9141a8ab09a1 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06cc2734-d903-4553-956d-dbfcdbbf5371 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ab807f-1c88-4b7d-a8ea-0832ef65912c · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0efc949a-7bae-46ad-8973-be955f65c7d2 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models A diagram is worth a dozen images
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a46e26cb-fe04-4b59-869d-dffdb9d6281a · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Learn to explain: Mul- timodal reasoning via thought chains for science question answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1fbbf652-46d6-4fd8-b1d8-d0fa655c23b0 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653df0b8-edd1-4d60-9a2d-b005808c9712 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Docvqa: A dataset for vqa on document images
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef3bfa21-e413-4e3e-8aea-086b37dfa983 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b86e8a-8257-437f-a287-5cc42410321a · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Towards vqa models that can read
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9ec0c7f-c951-41dc-82cd-e9054f08ef5b · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Grok-1.5 vision preview
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19656f93-4ef4-48a0-ae14-7711b892958c · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Gaussian Error Linear Units (GELUs)
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfdaa309-176d-42e2-a61c-75346a002ea2 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Mpnet: Masked 34 and permuted pre-training for language understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6965958c-569f-4dcd-ac64-8d22f287c9e9 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ef5fe6bd-c25c-41eb-b6bb-58a347ec8ae2 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Decoupled Weight Decay Regularization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b03c9c9-6175-447d-b0a2-a922ed74190e · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Vizwiz grand challenge: Answering visual questions from blind people
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 237d7dad-4893-422d-b544-c4588fb984f3 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f9fbc8-6894-4604-883a-f29cc4ee3007 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 08e6b8ce-f8e3-4346-be79-123858363622 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e17f659-684c-4700-acfd-ede212d60078 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Obelics: An open web-scale fil- tered dataset of interleaved image-text documents
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b4dce60f-e2f2-4024-b239-f0c80c3ef830 · outbound
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.