Pith. sign in

Paper Citation Record · LEDGER

Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2311.18021.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18021 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:14:08.070786Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:40:56.036122Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 805c69e3-cfec-4ccb-b641-a27539f7c170 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:08.070786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:08.070786Z digest=sha256:43b978e5a19e3bc5ade8099c5453803d998e7b6494b9e97401c501f3ae1c6e8e

Observation 0a990664-95bc-4ed7-9107-548f6eb72e66 · inbound

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models cites this paper.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.347857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.347857Z digest=sha256:1ab8c686c89c3b03bade2080bebf85b4329cbd89ed8e772591fc16e769c5d714

Observation d0efde7c-62a4-4f48-934c-376553f6b27a · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.637147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.637147Z digest=sha256:5c8f93ce67d3cb4feefe3eaef37ec78b5ab0f1929308189986e425ff6c863070

Observation f7798c12-b2f3-499a-afef-f652dc53e1b7 · inbound

Online In-Context Distillation for Low-Resource Vision Language Models cites this paper.

Online In-Context Distillation for Low-Resource Vision Language Models Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:40:56.038621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:36:38.735914Z digest=sha256:3f4496251789fafe230d5ce43c1f189e5ad053dda24c677220c659ab7d2a5750

Observation 488fc930-8120-4fcb-aefc-b2555ab85a12 · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.206259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:79f3d32e9f5ce57ace7f6df5d76ba9b6a00ff503bd29a5bca1192b93b61f1229