Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:35.934009Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 4 inbound Pith citation observations for arXiv:2507.15807.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:35.934009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T07:15:11.031811Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T20:47:22.870706Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e30b601e-a2b9-431a-aed0-d4e28defce69 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4777390-faa6-4602-84eb-a7e7ce3d8d9f · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Each task is designed with adjustable difficulty levels, such as more diverse con- cepts in novel concept binding, more complex visual patterns in pattern interpretation, etc
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5ea653a-54bd-4e0a-b88c-e7c1f2c6168f · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context A Survey on In-context Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efe63c4b-2f37-4df8-80bc-62a459a46624 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context In-context learning enables multimodal large language models to classify cancer pathology images
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e94f6ba-10e7-4fbf-a878-8bbf6543c081 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context In-context learning enables multimodal large language models to classify cancer pathology images
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81541ab2-1b20-4b79-8230-ef2d6bf3d74c · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Many-Shot In-Context Learning in Multimodal Foundation Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f949f0e5-46a8-445c-be4f-4a7c575cc8a6 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Many-Shot In-Context Learning in Multimodal Foundation Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3148e369-a2de-4a8f-aab3-1ace73c11c2d · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context MIBench: Evaluating Multimodal Large Language Models over Multiple Images
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9c221e-66ce-4533-8436-62b57fed643d · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context A survey on lora of large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97e71c88-e3a7-4d53-8d52-d14e98fc31ea · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context GPT-4o System Card
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e5bc96-48c3-4d02-83a9-7a99df77082b · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context GPT-4o System Card
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f34ef8a-b5f5-4b23-88f1-b4bde65a90f3 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb3b7ee-5dff-4e22-abba-519bf33b5195 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf365a9e-c200-4616-a198-f5ed12c00cd5 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7128cb3a-d652-462f-9b52-9734428ecd8b · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Learning to Retrieve In-Context Examples for Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535cdd76-764d-4c93-90c2-c5102b1a572a · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d74c232-7be3-496e-8ba4-d34d5a0e7702 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Low-rank adaptation for foundation models: A comprehensive review
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cec880b-4fb4-4243-b9f4-8972be25ee3b · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf00b6b2-469d-4226-94e6-843f3fb610e1 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f304a5c6-2cb3-4968-9053-1d34374fe6bb · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993d4abe-3f16-40b6-9fbd-24509c847c6e · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72745e32-46c3-46c6-a753-4f3062bbbf79 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Suppose we introduce an attention reallocation factor to the softmax operation by defining F := diag(f) ∈ RL×L, where f ∈ RL is a vector of learnable factors
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 76b1a58c-90ab-4a17-9672-254d8e38bfed · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bee96674-cf52-4d2f-ad4a-a0b00a25105c · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context As shown in the figure, within the range of a few hundred parameters, different configurations have no significant difference in the impact on the final performance
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97a50230-23ed-4cd0-a526-df69a1fde58d · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84e408d6-aba9-4656-aff9-84ee0bb82e34 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Advanced Multimodal Deep Learning Architecture for Image-Text Matching
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 712a4f79-f2dc-43e6-ab4e-2a170e58783a · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0efde7c-62a4-4f48-934c-376553f6b27a · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ca7f5c-4333-4f1f-ac79-bab58917cdb4 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1fab29-7ff0-4d7d-800c-e27fe95aa35b · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Towards Multimodal In-Context Learning for Vision & Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6621e578-bf53-499a-ad72-264555676780 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context LoRA: Low-Rank Adaptation of Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b3e3fd-2f3f-4d75-a503-2094b83ceb13 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Language models are few-shot learners
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56635091-ef64-4704-98cf-59e530b62c46 · outbound
True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228ec975-998a-422c-89b7-0bb4354606c1 · inbound
Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers True Multimodal In-Context Learning Needs Attention to the Visual Context
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49fda585-c582-4230-9647-634ebf07c115 · inbound
Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks True Multimodal In-Context Learning Needs Attention to the Visual Context
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 694c1bcf-b7dd-4af0-bf36-3dc8c2ea6832 · inbound
Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning True Multimodal In-Context Learning Needs Attention to the Visual Context
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f3145f01-dd25-4afc-8ca7-d6629c2f08fe · inbound
Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems True Multimodal In-Context Learning Needs Attention to the Visual Context
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.