Pith. sign in

Paper Citation Record · LEDGER

True Multimodal In-Context Learning Needs Attention to the Visual Context

As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 4 inbound Pith citation observations for arXiv:2507.15807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15807 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:35.934009Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T07:15:11.031811Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T20:47:22.870706Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e30b601e-a2b9-431a-aed0-d4e28defce69 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

True Multimodal In-Context Learning Needs Attention to the Visual Context Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.466845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.466845Z digest=sha256:c83bcde79608692c63c594a5919e226f810984c0d40f523e8dfccf8855642fc0

Observation e4777390-faa6-4602-84eb-a7e7ce3d8d9f · outbound

This paper cites Each task is designed with adjustable difficulty levels, such as more diverse con- cepts in novel concept binding, more complex visual patterns in pattern interpretation, etc.

True Multimodal In-Context Learning Needs Attention to the Visual Context Each task is designed with adjustable difficulty levels, such as more diverse con- cepts in novel concept binding, more complex visual patterns in pattern interpretation, etc

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:29:36.680644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:29:35.671431Z digest=sha256:958d514e8c0a3bce8f3d4ce90eb7ec5b56b23a271e19f762394ab257adfca103

Observation a5ea653a-54bd-4e0a-b88c-e7c1f2c6168f · outbound

This paper cites A Survey on In-context Learning.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Survey on In-context Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.727729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.727729Z digest=sha256:e48c915830681451024a5276bdedad52d339f3c0bddc352ed1a1b7269d7271c9

Observation efe63c4b-2f37-4df8-80bc-62a459a46624 · outbound

This paper cites In-context learning enables multimodal large language models to classify cancer pathology images.

True Multimodal In-Context Learning Needs Attention to the Visual Context In-context learning enables multimodal large language models to classify cancer pathology images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.873679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.873679Z digest=sha256:0810bcfb84bdd888c3787be29e23d71365c88e13aa7a49fb5e9e3af65327c3df

Observation 7e94f6ba-10e7-4fbf-a878-8bbf6543c081 · outbound

This paper cites In-context learning enables multimodal large language models to classify cancer pathology images.

True Multimodal In-Context Learning Needs Attention to the Visual Context In-context learning enables multimodal large language models to classify cancer pathology images

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:29:36.196407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:29:33.947604Z digest=sha256:d474c3b2a7e1ccb3b21d0d6cfa955410f9f0455efe0d7a55eedd862881075133

Observation 81541ab2-1b20-4b79-8230-ef2d6bf3d74c · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.320527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.320527Z digest=sha256:361091f2101b4502bf4b5e6c587f8c2eca95d46259706e85cd2841fdf22177f4

Observation f949f0e5-46a8-445c-be4f-4a7c575cc8a6 · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.377976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.377976Z digest=sha256:1e7b8a2e85c187516dda76e490a83ca117e6f0138da9567c59bfce85e86a8886

Observation 3148e369-a2de-4a8f-aab3-1ace73c11c2d · outbound

This paper cites MIBench: Evaluating Multimodal Large Language Models over Multiple Images.

True Multimodal In-Context Learning Needs Attention to the Visual Context MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.466045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.466045Z digest=sha256:243fa4ac55864391150d28bb6a6e2bf61d199215dc47468e36280e89c3bd3215

Observation dc9c221e-66ce-4533-8436-62b57fed643d · outbound

This paper cites A survey on lora of large language models.

True Multimodal In-Context Learning Needs Attention to the Visual Context A survey on lora of large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:29:36.710394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:29:34.538808Z digest=sha256:62e1a2853610e2d2704a991d1f0c236f4e6ce5fe91878e3ea74cccae0b52d218

Observation 97e71c88-e3a7-4d53-8d52-d14e98fc31ea · outbound

This paper cites GPT-4o System Card.

True Multimodal In-Context Learning Needs Attention to the Visual Context GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.609047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.609047Z digest=sha256:4fd6a6a7a781fee8d97e6be4bb3a7ec7882d1b7264771535b54de50a921642f1

Observation c9e5bc96-48c3-4d02-83a9-7a99df77082b · outbound

This paper cites GPT-4o System Card.

True Multimodal In-Context Learning Needs Attention to the Visual Context GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.684005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.684005Z digest=sha256:828bf4ff47f436babc4c78ab4afad4a3779364f6f88f5368243bdb85af3be64b

Observation 9f34ef8a-b5f5-4b23-88f1-b4bde65a90f3 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

True Multimodal In-Context Learning Needs Attention to the Visual Context What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.782618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.782618Z digest=sha256:f7aea47376aee7342e2ede15acd47709a7758c4bae95da5ac39ffb1dc656ab37

Observation 4bb3b7ee-5dff-4e22-abba-519bf33b5195 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.853072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.853072Z digest=sha256:89fdaaeafb192a339770522089628cdce9e7dd5901521da3179f2c71204af410

Observation cf365a9e-c200-4616-a198-f5ed12c00cd5 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.925120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.925120Z digest=sha256:0e33feeec88752092970c9d23bd88dd7e6e6ed527d4cf3d82a655f69e848fe4f

Observation 7128cb3a-d652-462f-9b52-9734428ecd8b · outbound

This paper cites Learning to Retrieve In-Context Examples for Large Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Learning to Retrieve In-Context Examples for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.040526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.040526Z digest=sha256:ed4517851fa2ddd6e054cfb6a18aa384d90ef3531969ebae76aae8924ced4583

Observation 535cdd76-764d-4c93-90c2-c5102b1a572a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

True Multimodal In-Context Learning Needs Attention to the Visual Context Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.096401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.096401Z digest=sha256:481b5eb59efbaf16542f130fa766192eb5ef621fa6aaafd512fc7c70b2b371e7

Observation 8d74c232-7be3-496e-8ba4-d34d5a0e7702 · outbound

This paper cites Low-rank adaptation for foundation models: A comprehensive review.

True Multimodal In-Context Learning Needs Attention to the Visual Context Low-rank adaptation for foundation models: A comprehensive review

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.167322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.167322Z digest=sha256:1ba7066afd91bef14c3baf95745d82621d45ca9d7af79ee53e38e66754fefd97

Observation 6cec880b-4fb4-4243-b9f4-8972be25ee3b · outbound

This paper cites In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation.

True Multimodal In-Context Learning Needs Attention to the Visual Context In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.243089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.243089Z digest=sha256:cecbe410145efa3aae75108db035f476890f0383a108bcfea7d1ccf15097053e

Observation cf00b6b2-469d-4226-94e6-843f3fb610e1 · outbound

This paper cites In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation.

True Multimodal In-Context Learning Needs Attention to the Visual Context In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.291674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.291674Z digest=sha256:097d9db523f5834b37a36c5458c91e5298beb62d47cfaf36e618192dad3a56c9

Observation f304a5c6-2cb3-4968-9053-1d34374fe6bb · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

True Multimodal In-Context Learning Needs Attention to the Visual Context MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.387308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.387308Z digest=sha256:c060781545f71eb2c668e2b9e32d1979b05299ef517ddc71a75840a166a946dc

Observation 993d4abe-3f16-40b6-9fbd-24509c847c6e · outbound

This paper cites VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning.

True Multimodal In-Context Learning Needs Attention to the Visual Context VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.474334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.474334Z digest=sha256:1f968baae074e50c310ba8d3f7d13431052cb12a860777d9fcb46c26d60928e4

Observation 72745e32-46c3-46c6-a753-4f3062bbbf79 · outbound

This paper cites Suppose we introduce an attention reallocation factor to the softmax operation by defining F := diag(f) ∈ RL×L, where f ∈ RL is a vector of learnable factors.

True Multimodal In-Context Learning Needs Attention to the Visual Context Suppose we introduce an attention reallocation factor to the softmax operation by defining F := diag(f) ∈ RL×L, where f ∈ RL is a vector of learnable factors

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:29:36.695880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:29:35.565985Z digest=sha256:42021a96e2a27c881d7162fa7f0f986a26614d44c0c1b47ea1939564c3845dd5

Observation 76b1a58c-90ab-4a17-9672-254d8e38bfed · outbound

This paper cites an unresolved cited work.

True Multimodal In-Context Learning Needs Attention to the Visual Context Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:29:36.665811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:29:35.793693Z digest=sha256:7237378c5c2318c3a7145e825e672208fbf90adba705cd3192eb792ad0a3ebe4

Observation bee96674-cf52-4d2f-ad4a-a0b00a25105c · outbound

This paper cites As shown in the figure, within the range of a few hundred parameters, different configurations have no significant difference in the impact on the final performance.

True Multimodal In-Context Learning Needs Attention to the Visual Context As shown in the figure, within the range of a few hundred parameters, different configurations have no significant difference in the impact on the final performance

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:29:36.652176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:29:35.934009Z digest=sha256:2256892b311ff68f15a9637a82701cc41845380300007ec05e740facad183a60

Observation 97a50230-23ed-4cd0-a526-df69a1fde58d · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.523551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.523551Z digest=sha256:cf238096d53cc13e4ae22ca44b20dde4f1dd2e9248eee3cb538576e9a8b93a31

Observation 84e408d6-aba9-4656-aff9-84ee0bb82e34 · outbound

This paper cites Advanced Multimodal Deep Learning Architecture for Image-Text Matching.

True Multimodal In-Context Learning Needs Attention to the Visual Context Advanced Multimodal Deep Learning Architecture for Image-Text Matching

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:29:36.122777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:29:34.962587Z digest=sha256:5f96a27401324f56ec6d5c99009db63620353f9e2ab56c52b9148a5ba1b408c0

Observation 712a4f79-f2dc-43e6-ab4e-2a170e58783a · outbound

This paper cites SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization.

True Multimodal In-Context Learning Needs Attention to the Visual Context SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.240882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.240882Z digest=sha256:30a49b1f8eea0f455e2d452e9db62b75948db52dfd6a4f58202539539b769093

Observation d0efde7c-62a4-4f48-934c-376553f6b27a · outbound

This paper cites Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?.

True Multimodal In-Context Learning Needs Attention to the Visual Context Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.637147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.637147Z digest=sha256:5c8f93ce67d3cb4feefe3eaef37ec78b5ab0f1929308189986e425ff6c863070

Observation f5ca7f5c-4333-4f1f-ac79-bab58917cdb4 · outbound

This paper cites Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.173300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.173300Z digest=sha256:42ab7ab7f8960aa153b5fbeee9ee2dfba200d9bef00d6499217e8bd70744aea6

Observation da1fab29-7ff0-4d7d-800c-e27fe95aa35b · outbound

This paper cites Towards Multimodal In-Context Learning for Vision & Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Towards Multimodal In-Context Learning for Vision & Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.800157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.800157Z digest=sha256:8068ba9dd69813ff4b9fbb7f21a0aad4add4d9ae6367604e4ba7f29bd442780d

Observation 6621e578-bf53-499a-ad72-264555676780 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context LoRA: Low-Rank Adaptation of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.121756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.121756Z digest=sha256:1c312d0b82cd07a9e65f8957832f14c0cbe37f589e0fb55260c53522236a2e34

Observation 95b3e3fd-2f3f-4d75-a503-2094b83ceb13 · outbound

This paper cites Language models are few-shot learners.

True Multimodal In-Context Learning Needs Attention to the Visual Context Language models are few-shot learners

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.589687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.589687Z digest=sha256:112d4c61ce649861bf0d552ba0082eb1061ea71e91c4804e0ce0f820d2874e6d

Observation 56635091-ef64-4704-98cf-59e530b62c46 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.010676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.010676Z digest=sha256:829d305abe48fe19e23ec4d5472edf20cd348651e03468bc32b95a82bc6b76ef

Pith citing papers

Observation 228ec975-998a-422c-89b7-0bb4354606c1 · inbound

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers cites this paper.

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers True Multimodal In-Context Learning Needs Attention to the Visual Context

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T07:15:11.031811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:15:11.031811Z digest=sha256:f54d864342440ce3c3733fd909072ab03cd8540805cbfab0722b0b99fd63e5e9

Observation 49fda585-c582-4230-9647-634ebf07c115 · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks True Multimodal In-Context Learning Needs Attention to the Visual Context

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.210983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:903e2276b1ed70f47d2a2fb9f1badf79c3d0c1faf32015df3c16b8936225c11b

Observation 694c1bcf-b7dd-4af0-bf36-3dc8c2ea6832 · inbound

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning cites this paper.

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning True Multimodal In-Context Learning Needs Attention to the Visual Context

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.182673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:21:30.235583Z digest=sha256:20f397468759a5a56b94913b6f999ef342a6975612162f6c7852eb3851d89bd2

Observation f3145f01-dd25-4afc-8ca7-d6629c2f08fe · inbound

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems cites this paper.

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems True Multimodal In-Context Learning Needs Attention to the Visual Context

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.872136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T20:11:02.445626Z digest=sha256:13621c7242484649c51b65b6de498c7d24a0eee2561f48985220ddd1f5d572a2