Pith. sign in

Paper Citation Record · LEDGER

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 5 inbound Pith citation observations for arXiv:2506.07936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07936 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:30.497174Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:43:16.986690Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b5eba608-dd3f-43d8-809a-2737621f1a43 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.326469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.326469Z digest=sha256:6fc150d5e61677e0332e0879baf678faf4c8880ee1bd91a1c491066982245021

Observation 6ac019dc-ebe7-4ccd-93b3-2005b827e975 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.330578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.330578Z digest=sha256:394e8ae2e2f7a47cf08dc60dc761d7d3e57f0dc9d4611a024f3699b8c127324c

Observation 0206e29a-924e-436e-8dc7-d959a0587dfe · outbound

This paper cites Qwen2.5-VL Technical Report.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.333944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.333944Z digest=sha256:6ec1ee66e7febb6dec1b289dba308ff857a6dd529dff543369cd863fb2d584b0

Observation 66e436cd-5e95-4931-ae50-c3bd46f26429 · outbound

This paper cites What makes multimodal in-context learning work? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1539–1550, 2024.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What makes multimodal in-context learning work? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1539–1550, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.337467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.337467Z digest=sha256:d802b8addaac1aa17b161d225d00fac75a9540f79f6230638b9c0243c0745553

Observation 1b60879c-f6ec-499a-a516-da1b4fdfefe6 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.340707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.340707Z digest=sha256:ecf6a037c00ddf5527c475385c21437da34fccd0f6a1ec3411e009a97c149067

Observation 5cfa213e-4e85-4600-8e5a-d68347522cd7 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.344139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.344139Z digest=sha256:156959a0580bda81ebf420cedad6a68c06df167ec75a86bc7dc7ce4e871cf7ba

Observation 0a990664-95bc-4ed7-9107-548f6eb72e66 · outbound

This paper cites Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.347857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.347857Z digest=sha256:0f52080b5196e0d4ed3087af7e8aa1d9b1bc40e78e493a1839ca23cba726514a

Observation d7134849-2ec4-4fe9-9076-ec3d3ab643e4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.352098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.352098Z digest=sha256:d091b1f0c86d76d287642dad6d954eb69beba287e555e19928277b2ff5cbe67d

Observation b3be597a-86b0-4857-a4ed-2e67d0c0c170 · outbound

This paper cites A Survey on In-context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models A Survey on In-context Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.355259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.355259Z digest=sha256:0f8d55252946636bbc9f70d251696a8344004024dca394562bcba1fb7c1bb947

Observation 3124b6bc-7b85-49ed-8d4a-6cc563b1ffcb · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.358643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.358643Z digest=sha256:1d60308573a95f6026713a630d0bf05093816f62f1fb2ef9ad3692e114fbef9b

Observation 79eb36fa-ddc5-4e7c-9127-09243a67138f · outbound

This paper cites Interleaved-Modal Chain-of-Thought.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Interleaved-Modal Chain-of-Thought

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.361398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.361398Z digest=sha256:6bb8f112eb35fcbd2a583fb99f2d2e0b2df798285bb628c2a034497ba4bd5132

Observation caf709a0-2700-4a83-a67d-e1ac8b44295c · outbound

This paper cites Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.364635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.364635Z digest=sha256:cb922ad28f5db2e10819ec54da6897142870eeb634a27afc9b7c58a12835ae63

Observation 58541b15-d921-4064-b89d-6dee79b1e6b8 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.367946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.367946Z digest=sha256:b8721fd1de362a44b28c51fca1efd85aee2c4f1bac6def24f8d2bfba965f20fa

Observation cac6d4c7-64c9-4415-9f84-4ca332bb8d68 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.371674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.371674Z digest=sha256:1ecb946b3c7958db5b0433fe49377cc0ab7aea58580e0efee7522bcc1d12dc5b

Observation 725bc155-8341-46d3-a66d-1a6395d71782 · outbound

This paper cites Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.375814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.375814Z digest=sha256:4cefeefdd113ec3542661ead2d29de10b1f5389f8872ea75601636cc4c87dda4

Observation be9d5f8a-08dc-47bc-a483-72007514d435 · outbound

This paper cites What matters when building vision-language models?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What matters when building vision-language models?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.379020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.379020Z digest=sha256:a2b3d704def3fbb94a9ee8e533624a550942af765ac18fbeeca465b178f2db4f

Observation a003c4e3-e160-4134-952d-a5187c342958 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.381931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.381931Z digest=sha256:e80abd096af148f7c08adb41dd2d08ddfc0d62bde5d033fe999a8fe3c38fd742

Observation c9adec86-6858-4838-8f48-a51fcd794183 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.385062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.385062Z digest=sha256:666feba89d90633f1852999deca12c45346c19389976c99004396e9047d6a99b

Observation ee0fcb41-a038-4742-940b-4121a9091338 · outbound

This paper cites Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.388110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.388110Z digest=sha256:854a5da34f32c38ec6a0d740e4eb8e15bb7eda60bac3cf0a23c7746224994eff

Observation 6903f7b2-2961-4c17-8107-a935e8baa3e8 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.https: //ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.https: //ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.049043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.391394Z digest=sha256:c76d72fb894bc046a79690acd624cefc67647846f845d4d93518aae5a3e38052

Observation 1a023700-31db-4bd3-be62-7c1147e17792 · outbound

This paper cites Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.394417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.394417Z digest=sha256:e0f4e890deef8c0cf62aeef58a0227a66ca041a28343cb43d0e74ce442154c52

Observation f8ae2fe6-4446-47d9-9941-f825c3bfb4e9 · outbound

This paper cites Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.397479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.397479Z digest=sha256:871a476df2e42553ae4cc0a5efb56ac9ce5b42bdddefa24a860711cf991ebeb1

Observation 9d60bd3c-8794-4bbd-9635-e66877b7d1e4 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.400396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.400396Z digest=sha256:d7b0f3bbb74eb4c85935d8919861836283ad9277fa82b0020355a3fc887b0784

Observation 0d358e1e-a868-4533-be8e-b2a645030683 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.403345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.403345Z digest=sha256:f91cc15a5f093ab4471551f693b9e2704abc4aac02e66bb98a971d688ff2529e

Observation cdd8f034-939e-43fc-bb23-0587c90b662e · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.406334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.406334Z digest=sha256:b8130daca9cc10ea21c015b425601d97f7b7db85e49c3cc5ab2a488cdaf49275

Observation 905500d6-2fba-4677-9eb9-ce6a287ce378 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.409388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.409388Z digest=sha256:e9f9554e08c9fa91276d9f1ee193bf65a1d2c6d29733696218ef9f07156f9e59

Observation c79db2fd-71d1-418f-af04-e7c53c07e56f · outbound

This paper cites Towards vqa models that can read.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.040011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.412385Z digest=sha256:1db20633e4a781d92b07a1f98f44e5897b6a50031e92c580c38ecd9c4e44b57a

Observation 51a84144-6660-4ddd-8a42-896548d3a749 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.415272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.415272Z digest=sha256:5352266087980073dcdf44960b32d6be4ba878589aee1faf5348644117d44530

Observation 9fef8f3e-b170-4db1-9add-e81f420a39da · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.418218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.418218Z digest=sha256:064dd6921696cd7a4b0cc9c74546c744024a3a5e62d3dbec30714129815c5481

Observation ed83cb5c-4e55-47a7-a25b-9fdef8c40abb · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.421112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.421112Z digest=sha256:2c186ddffe9e0dc288eb7c01283c9272708cc36d17cc8aa514732c14aa7220d1

Observation 269f6903-dc82-4e4e-b990-fe6d1046415c · outbound

This paper cites Emergent Abilities of Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Emergent Abilities of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.424057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.424057Z digest=sha256:21be95619e6ee79f2d6cbf4f3069a4110eecf4b4a955ad50f81025e6a46eded6

Observation 0db51feb-290f-4c93-b4ca-b6f292731f52 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.427120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.427120Z digest=sha256:079290053e90f51709cf45b07d02a6cafef9c0dbe4f71bf9fbf770cd0b50536c

Observation 0aad04a6-6109-44b6-be93-8f56153d5371 · outbound

This paper cites Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.430572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.430572Z digest=sha256:cbeae865d4f819670ece5fbaf027e22a51d2ec281697da61c7cbe23e0db1b99a

Observation 736cf377-7b4e-4e2b-8d0a-bb9eff84a94c · outbound

This paper cites Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:25:30.646131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.433663Z digest=sha256:8ccde4e6745d45de7779e3699a941592f3b18f1d7549b378ec313c6fccf132b1

Observation 0a3eec10-ee2b-4186-89bb-5b762e7a11a6 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.436700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.436700Z digest=sha256:08a2858ca1489ded561d7df2895c4c7dc8438018846d7b07681e701b8e65bad2

Observation ef44611e-ed1c-4af7-8c56-8c74eb41f55c · outbound

This paper cites From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.439676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.439676Z digest=sha256:d74f4bf94a966963fddb555e114531cced29ca92df3e68f38b609ea46a7f4351

Observation 56d2d879-eda2-45fa-a7d1-d766326c8e85 · outbound

This paper cites Formal Mathematical Reasoning: A New Frontier in AI.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Formal Mathematical Reasoning: A New Frontier in AI

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.442609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.442609Z digest=sha256:44149d598f0c55f5cdca82e8c330144c7986bfdd10f9c593cdec604f250332cf

Observation 35657318-eab6-4585-a89d-41f3582b7ffb · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.031434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.445696Z digest=sha256:4973bbe53038a1de4e631237701179d5bb83f65fb5d2aa3e7df3acd65d1826f6

Observation 3d2ce3f7-cb96-49b9-a637-0ebc28f48b68 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Automatic Chain of Thought Prompting in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.448415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.448415Z digest=sha256:83a813f704b35b0ffd1e6a74242dda7b3cdb77001a919d26000b71f94c0b4010

Observation cf7df1ba-8ce8-49cd-91ef-48a135e90ccb · outbound

This paper cites Wong, and Simon See.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Wong, and Simon See

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.451593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.451593Z digest=sha256:36536c7a770089a3641a2ddfa9e0b288bd1ffae0a148ed3155b878e7c13e1d29

Observation ae870c12-76c4-4867-99aa-71454747ca65 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.454384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.454384Z digest=sha256:24057aca066a3e782c9932a2f185ec9f1f3398f66d84b2242f106a111cd05099

Observation 964eca0e-c530-40d4-8056-76e070db338b · outbound

This paper cites VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.457496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.457496Z digest=sha256:18b71524016f783e8f7c704b3c398201e1f3ddff0fa02bc5bafb6f4c3f6b19b6

Observation bd652441-f733-40b4-bdb2-c4b0485883a4 · outbound

This paper cites Can In-context Learners Learn a Reasoning Concept from Demonstrations?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can In-context Learners Learn a Reasoning Concept from Demonstrations?

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:25:30.527381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.460503Z digest=sha256:1ad8e8fc465b5fc04b4d932716a7d7518274051842c8ddd7eea1c661dab0e449

Observation 4a36758c-10e8-4e96-bdba-3a2e18ee3d68 · outbound

This paper cites Density is calculated as mass/volume.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Density is calculated as mass/volume

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.022770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.464287Z digest=sha256:f0bf081e22b3de3d586da4d06e340d9b26bf5bd199d3afba11c50f578e14c1ad

Observation 574fe031-2fd9-4b20-aa9b-b76a3abdb119 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:25:31.014289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.467233Z digest=sha256:597e70bfcecf44bac3451df09abba2681f6e736bd9a91f6a6ca8c65f1c6c4241

Observation d43821c7-b720-4112-ab66-d3ff04106176 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:31.005482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.470376Z digest=sha256:d83e827d8598cc853194267a9d2f22cf9e7e6895ac114c66e765c149c849fb90

Observation 084a3b22-a8ff-479f-9e4c-88a6f6bbcafa · outbound

This paper cites Final answer:.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Final answer:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.996962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.473371Z digest=sha256:251680f06ff0668c932e58e35b733e8a8d9e1986ae6968ace6953da68f294182

Observation a5b37fc7-91ce-4e09-860f-d97eaaba693e · outbound

This paper cites In the graphs, we need to compare the export value (top graph) with the import value (bottom graph) for each country in the year 2020.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models In the graphs, we need to compare the export value (top graph) with the import value (bottom graph) for each country in the year 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.988293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.476297Z digest=sha256:1cc16571170a52d7905191c1e8258440e0b9d70edd2f0578419bcd28dd4179d6

Observation 99a3edf5-38dd-49e6-afcf-e42e0a996b9b · outbound

This paper cites Export > Import, so Country 3 has a surplus.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Export > Import, so Country 3 has a surplus

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.978120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.479128Z digest=sha256:d3299aad70353607113ed5c0fa5717eb7d56eab8f698b627d017fb6165ae6b3e

Observation dcf0a135-2d0a-411b-b19e-a5c855ae7879 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:30.969273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.482045Z digest=sha256:fec9d79faf1717d01a90a35784abf62554f88d4a2ca7317568a79ce811c6b814

Observation 2a9c27f2-c4f4-41c8-9844-8a492ad6ffd2 · outbound

This paper cites Final answer:.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Final answer:

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.959723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.484729Z digest=sha256:f3c17fbd46bd8733b430332764e31ba92c31187064e060ae9f86047b585be925

Observation 4c6fbc48-6bfe-4c30-bba2-6546a303e330 · outbound

This paper cites a² varies inversely with b³.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models a² varies inversely with b³

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.949945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.487895Z digest=sha256:919f614008734a58c91ff95ff4ec30731b514c4994978a51a5fba736fa0c352e

Observation 3e46e54a-673e-47f6-a7ba-0ae533126ade · outbound

This paper cites Therefore, a² = 7² = 49.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Therefore, a² = 7² = 49

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.941221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.491118Z digest=sha256:c5fc969cd777d8b31308bb51be781697ed8c55620732298aa62036e094beecf3

Observation e64c6f14-4372-42e7-b695-b4ca36a8aab3 · outbound

This paper cites When b = 6, a² = 1323 / 6³ = 1323 / 216 = 6.125.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models When b = 6, a² = 1323 / 6³ = 1323 / 216 = 6.125

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.932322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.494019Z digest=sha256:73fb6744d2b8e3b83ecf99bd4f5dc67cc328822832eba49e89aae5a081285ead

Observation 54bdac12-e0f8-42b8-92ef-2625aef72dec · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:30.923432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:25:30.497174Z digest=sha256:46e160483eed37deb30e3755f1f814a85aa68f8710a70260a7c533e59adbbc89

Pith citing papers

Observation adec5307-bff3-4d51-9fa7-21126ac3a9bb · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.986690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.986690Z digest=sha256:56d7a72e66fdb1b53a87aa217818c1d64cf8bf3363a6c7b6e61851c8ec56cbe6

Observation f5ca7f5c-4333-4f1f-ac79-bab58917cdb4 · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.173300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.173300Z digest=sha256:12a64218c9fe8a2b3bf28f6cac1f5f50e97e6081f1eca8fea430a64686c5b3ba

Observation 84d6614a-312c-4db4-b403-35f50e49d566 · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.448460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:8d01502c1feb0b7db8d3e70bf3a2ce1913ecd2e36a6bd0eb4dda8be8b4061480

Observation 957de704-b9a9-4c56-983f-e25012054cec · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.245354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:ccf4a3d3649af13be1c0d0ea795b2874982556d9c8b521092afb79fc5a87b371

Observation 1dae7721-237e-4948-8945-f6b27325da46 · inbound

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice cites this paper.

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:50.173877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:03:50.055081Z digest=sha256:f7e9de6be31418a1bae5fa7fdf3593287dd9666a8a759f74d6c673f72bda7078