Pith. sign in

Paper Citation Record · LEDGER

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

As of 9 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 5 inbound Pith citation observations for arXiv:2506.07936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07936 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:30.497174Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:43:16.986690Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b5eba608-dd3f-43d8-809a-2737621f1a43 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.326469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.326469Z digest=sha256:6fc150d5e61677e0332e0879baf678faf4c8880ee1bd91a1c491066982245021

Observation 6ac019dc-ebe7-4ccd-93b3-2005b827e975 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.330578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.330578Z digest=sha256:394e8ae2e2f7a47cf08dc60dc761d7d3e57f0dc9d4611a024f3699b8c127324c

Observation 0206e29a-924e-436e-8dc7-d959a0587dfe · outbound

This paper cites Qwen2.5-VL Technical Report.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.333944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.333944Z digest=sha256:6ec1ee66e7febb6dec1b289dba308ff857a6dd529dff543369cd863fb2d584b0

Observation 66e436cd-5e95-4931-ae50-c3bd46f26429 · outbound

This paper cites What makes multimodal in-context learning work? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1539–1550, 2024.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What makes multimodal in-context learning work? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1539–1550, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.337467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.337467Z digest=sha256:d802b8addaac1aa17b161d225d00fac75a9540f79f6230638b9c0243c0745553

Observation 1b60879c-f6ec-499a-a516-da1b4fdfefe6 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.340707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.340707Z digest=sha256:ecf6a037c00ddf5527c475385c21437da34fccd0f6a1ec3411e009a97c149067

Observation 5cfa213e-4e85-4600-8e5a-d68347522cd7 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.344139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.344139Z digest=sha256:f620878f5960d9cd11587103a2321e2af4e300822ba78578a35fa0b9b9683e68

Observation 0a990664-95bc-4ed7-9107-548f6eb72e66 · outbound

This paper cites Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.347857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.347857Z digest=sha256:1ab8c686c89c3b03bade2080bebf85b4329cbd89ed8e772591fc16e769c5d714

Observation d7134849-2ec4-4fe9-9076-ec3d3ab643e4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.352098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.352098Z digest=sha256:d091b1f0c86d76d287642dad6d954eb69beba287e555e19928277b2ff5cbe67d

Observation b3be597a-86b0-4857-a4ed-2e67d0c0c170 · outbound

This paper cites A Survey on In-context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models A Survey on In-context Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.355259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.355259Z digest=sha256:0f8d55252946636bbc9f70d251696a8344004024dca394562bcba1fb7c1bb947

Observation 3124b6bc-7b85-49ed-8d4a-6cc563b1ffcb · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.358643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.358643Z digest=sha256:1d60308573a95f6026713a630d0bf05093816f62f1fb2ef9ad3692e114fbef9b

Observation 79eb36fa-ddc5-4e7c-9127-09243a67138f · outbound

This paper cites Interleaved-Modal Chain-of-Thought.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Interleaved-Modal Chain-of-Thought

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.361398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.361398Z digest=sha256:66113fef32e77071e9b0f311e187dd5b56c4460d8e566121e0dc6592c5064a4a

Observation caf709a0-2700-4a83-a67d-e1ac8b44295c · outbound

This paper cites Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.364635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.364635Z digest=sha256:cb922ad28f5db2e10819ec54da6897142870eeb634a27afc9b7c58a12835ae63

Observation 58541b15-d921-4064-b89d-6dee79b1e6b8 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.367946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.367946Z digest=sha256:f93c4211c576c74d47bc658ba52ef750c949cfbdfc76d9fe8582488af555a176

Observation cac6d4c7-64c9-4415-9f84-4ca332bb8d68 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.371674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.371674Z digest=sha256:1ecb946b3c7958db5b0433fe49377cc0ab7aea58580e0efee7522bcc1d12dc5b

Observation 725bc155-8341-46d3-a66d-1a6395d71782 · outbound

This paper cites Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.375814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.375814Z digest=sha256:4cefeefdd113ec3542661ead2d29de10b1f5389f8872ea75601636cc4c87dda4

Observation be9d5f8a-08dc-47bc-a483-72007514d435 · outbound

This paper cites What matters when building vision-language models?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What matters when building vision-language models?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.379020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.379020Z digest=sha256:a2b3d704def3fbb94a9ee8e533624a550942af765ac18fbeeca465b178f2db4f

Observation a003c4e3-e160-4134-952d-a5187c342958 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.381931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.381931Z digest=sha256:e80abd096af148f7c08adb41dd2d08ddfc0d62bde5d033fe999a8fe3c38fd742

Observation c9adec86-6858-4838-8f48-a51fcd794183 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.385062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.385062Z digest=sha256:666feba89d90633f1852999deca12c45346c19389976c99004396e9047d6a99b

Observation ee0fcb41-a038-4742-940b-4121a9091338 · outbound

This paper cites Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.388110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.388110Z digest=sha256:854a5da34f32c38ec6a0d740e4eb8e15bb7eda60bac3cf0a23c7746224994eff

Observation 6903f7b2-2961-4c17-8107-a935e8baa3e8 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.https: //ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.https: //ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.049043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.391394Z digest=sha256:7aded9317487da37130e171148b0c0a07de853e12d33a371025d7d607b7c89cf

Observation 1a023700-31db-4bd3-be62-7c1147e17792 · outbound

This paper cites Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.394417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.394417Z digest=sha256:e007fc50bf7c00669ed2834d039e540922f278afea86044f3105dcef0526b6bf

Observation f8ae2fe6-4446-47d9-9941-f825c3bfb4e9 · outbound

This paper cites Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.397479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.397479Z digest=sha256:b367d1c2e0af3d5de11e5a41447baf1229c2f54437f417ef408398617dcc4e6a

Observation 9d60bd3c-8794-4bbd-9635-e66877b7d1e4 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.400396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.400396Z digest=sha256:d7b0f3bbb74eb4c85935d8919861836283ad9277fa82b0020355a3fc887b0784

Observation 0d358e1e-a868-4533-be8e-b2a645030683 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.403345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.403345Z digest=sha256:f91cc15a5f093ab4471551f693b9e2704abc4aac02e66bb98a971d688ff2529e

Observation cdd8f034-939e-43fc-bb23-0587c90b662e · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.406334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.406334Z digest=sha256:685bb7730fb95951af4c1623bf8790e8baa5becc068bbdc61ada732025b0cf8c

Observation 905500d6-2fba-4677-9eb9-ce6a287ce378 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.409388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.409388Z digest=sha256:e9f9554e08c9fa91276d9f1ee193bf65a1d2c6d29733696218ef9f07156f9e59

Observation c79db2fd-71d1-418f-af04-e7c53c07e56f · outbound

This paper cites Towards vqa models that can read.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.040011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.412385Z digest=sha256:323ee5fa98cc1a06de231c33196f67e452a5ac91eb224bd52e9eb7dbf9fa9e3d

Observation 51a84144-6660-4ddd-8a42-896548d3a749 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.415272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.415272Z digest=sha256:5352266087980073dcdf44960b32d6be4ba878589aee1faf5348644117d44530

Observation 9fef8f3e-b170-4db1-9add-e81f420a39da · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.418218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.418218Z digest=sha256:064dd6921696cd7a4b0cc9c74546c744024a3a5e62d3dbec30714129815c5481

Observation ed83cb5c-4e55-47a7-a25b-9fdef8c40abb · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.421112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.421112Z digest=sha256:2c186ddffe9e0dc288eb7c01283c9272708cc36d17cc8aa514732c14aa7220d1

Observation 269f6903-dc82-4e4e-b990-fe6d1046415c · outbound

This paper cites Emergent Abilities of Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Emergent Abilities of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.424057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.424057Z digest=sha256:21be95619e6ee79f2d6cbf4f3069a4110eecf4b4a955ad50f81025e6a46eded6

Observation 0db51feb-290f-4c93-b4ca-b6f292731f52 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.427120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.427120Z digest=sha256:079290053e90f51709cf45b07d02a6cafef9c0dbe4f71bf9fbf770cd0b50536c

Observation 0aad04a6-6109-44b6-be93-8f56153d5371 · outbound

This paper cites Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.430572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.430572Z digest=sha256:cbeae865d4f819670ece5fbaf027e22a51d2ec281697da61c7cbe23e0db1b99a

Observation 736cf377-7b4e-4e2b-8d0a-bb9eff84a94c · outbound

This paper cites Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:25:30.646131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.433663Z digest=sha256:8001896aa71759570c8e3208b73535fa6d3a1054bee9a4a445dbb53730ae3a45

Observation 0a3eec10-ee2b-4186-89bb-5b762e7a11a6 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.436700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.436700Z digest=sha256:08a2858ca1489ded561d7df2895c4c7dc8438018846d7b07681e701b8e65bad2

Observation ef44611e-ed1c-4af7-8c56-8c74eb41f55c · outbound

This paper cites From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.439676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.439676Z digest=sha256:c21f6ec5d88031a7bd9e855b04b9d84e5c8573394d70fa13c783c5fd0d585fb8

Observation 56d2d879-eda2-45fa-a7d1-d766326c8e85 · outbound

This paper cites Formal Mathematical Reasoning: A New Frontier in AI.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Formal Mathematical Reasoning: A New Frontier in AI

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.442609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.442609Z digest=sha256:bbb1e8d4d41b3b32503a04598f030cf1527bff9993b5c00799615829a5ab401a

Observation 35657318-eab6-4585-a89d-41f3582b7ffb · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.031434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.445696Z digest=sha256:523af9d6cb5785a06bd65429419cc7fa63c22ff35708b5b39377c6ddf5fe18a2

Observation 3d2ce3f7-cb96-49b9-a637-0ebc28f48b68 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Automatic Chain of Thought Prompting in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.448415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.448415Z digest=sha256:83a813f704b35b0ffd1e6a74242dda7b3cdb77001a919d26000b71f94c0b4010

Observation cf7df1ba-8ce8-49cd-91ef-48a135e90ccb · outbound

This paper cites Wong, and Simon See.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Wong, and Simon See

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.451593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.451593Z digest=sha256:36536c7a770089a3641a2ddfa9e0b288bd1ffae0a148ed3155b878e7c13e1d29

Observation ae870c12-76c4-4867-99aa-71454747ca65 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.454384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.454384Z digest=sha256:24057aca066a3e782c9932a2f185ec9f1f3398f66d84b2242f106a111cd05099

Observation 964eca0e-c530-40d4-8056-76e070db338b · outbound

This paper cites VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.457496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.457496Z digest=sha256:18b71524016f783e8f7c704b3c398201e1f3ddff0fa02bc5bafb6f4c3f6b19b6

Observation bd652441-f733-40b4-bdb2-c4b0485883a4 · outbound

This paper cites Can In-context Learners Learn a Reasoning Concept from Demonstrations?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can In-context Learners Learn a Reasoning Concept from Demonstrations?

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:25:30.527381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.460503Z digest=sha256:f439ce0a9e7117c9a068f5678639c9ac601295760946475cb951839ff862bc8f

Observation 4a36758c-10e8-4e96-bdba-3a2e18ee3d68 · outbound

This paper cites Density is calculated as mass/volume.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Density is calculated as mass/volume

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.022770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.464287Z digest=sha256:1799327581a19d8cf00091fd1917776527730b9913b443ddfa7108c9dd0e4a06

Observation 574fe031-2fd9-4b20-aa9b-b76a3abdb119 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:25:31.014289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.467233Z digest=sha256:53f016742b00bfd25af51867b110b9de8fcbf97a2a3b44e7ec3d27fbf392d39d

Observation d43821c7-b720-4112-ab66-d3ff04106176 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:31.005482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.470376Z digest=sha256:09a46ce60fe53f8623185e73c6acb2a06b391db47277c4c7cbbecf02750ec9b5

Observation 084a3b22-a8ff-479f-9e4c-88a6f6bbcafa · outbound

This paper cites Final answer:.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Final answer:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.996962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.473371Z digest=sha256:592babb8ca489117514f0db295d1f69dd74ed0bbb53159d59e4f2445c00991d4

Observation a5b37fc7-91ce-4e09-860f-d97eaaba693e · outbound

This paper cites In the graphs, we need to compare the export value (top graph) with the import value (bottom graph) for each country in the year 2020.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models In the graphs, we need to compare the export value (top graph) with the import value (bottom graph) for each country in the year 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.988293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.476297Z digest=sha256:e1239696fd4f9db342be3e3addeb96db7256945a46c6af8cb46462dadcf6c6ec

Observation 99a3edf5-38dd-49e6-afcf-e42e0a996b9b · outbound

This paper cites Export > Import, so Country 3 has a surplus.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Export > Import, so Country 3 has a surplus

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.978120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.479128Z digest=sha256:5dd0804eae673a47cf103ad5ad0ac8f37a03dc72982ccd7a3d8f9975e7fe2945

Observation dcf0a135-2d0a-411b-b19e-a5c855ae7879 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:30.969273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.482045Z digest=sha256:269152891eeeefb9c8e3e0ff41be0d4a67fd8fe814757502ba3badc1c3b5d305

Observation 2a9c27f2-c4f4-41c8-9844-8a492ad6ffd2 · outbound

This paper cites Final answer:.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Final answer:

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.959723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.484729Z digest=sha256:e9e2af9c27642a87ad6ccc8c618ea62d43ed3db37689eadbaea66e31c2154008

Observation 4c6fbc48-6bfe-4c30-bba2-6546a303e330 · outbound

This paper cites a² varies inversely with b³.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models a² varies inversely with b³

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.949945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.487895Z digest=sha256:fb30a9962f67aa1bbb18d4d6902047f556510bd41f56f30f099027ecaf850692

Observation 3e46e54a-673e-47f6-a7ba-0ae533126ade · outbound

This paper cites Therefore, a² = 7² = 49.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Therefore, a² = 7² = 49

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.941221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.491118Z digest=sha256:67d82c9bae0ac9a1ad7f2062b188a13d8bf34cda09ca1b5d67207c810269c1e9

Observation e64c6f14-4372-42e7-b695-b4ca36a8aab3 · outbound

This paper cites When b = 6, a² = 1323 / 6³ = 1323 / 216 = 6.125.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models When b = 6, a² = 1323 / 6³ = 1323 / 216 = 6.125

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.932322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.494019Z digest=sha256:d89a749af3bf4ca95ca288b9ad16538877c023cdc9b3baaad89defd140a6c008

Observation 54bdac12-e0f8-42b8-92ef-2625aef72dec · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:30.923432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:25:30.497174Z digest=sha256:ed2351314b5c3e64d50f6c932abbfea8d70c03deeae87b7d42b811fe9ba3b624

Pith citing papers

Observation adec5307-bff3-4d51-9fa7-21126ac3a9bb · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.986690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.986690Z digest=sha256:56d7a72e66fdb1b53a87aa217818c1d64cf8bf3363a6c7b6e61851c8ec56cbe6

Observation f5ca7f5c-4333-4f1f-ac79-bab58917cdb4 · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.173300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.173300Z digest=sha256:42ab7ab7f8960aa153b5fbeee9ee2dfba200d9bef00d6499217e8bd70744aea6

Observation 84d6614a-312c-4db4-b403-35f50e49d566 · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.448460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:82a38b211655e9332f271d57e3717308c3d27d6c164b84e71f86b259c88ccb29

Observation 957de704-b9a9-4c56-983f-e25012054cec · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.245354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:42b8b0feadce1d16f5c87a23aaeaa2560717230ff49a351a0bf1b00ffac2100c

Observation 1dae7721-237e-4948-8945-f6b27325da46 · inbound

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice cites this paper.

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:50.173877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:03:50.055081Z digest=sha256:d4754b8a002ee0782f7eb6d5dc0fa41697f7c330ab03b8d7bcdb82929c95f664