Pith. sign in

Paper Citation Record · LEDGER

Towards Multimodal In-Context Learning for Vision & Language Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2403.12736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.12736 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:18:30.950663Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:45:28.217864Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 47077c04-7d8b-41be-b03f-26ad4c5d2011 · inbound

Error-driven Data-efficient Large Multimodal Model Tuning cites this paper.

Error-driven Data-efficient Large Multimodal Model Tuning Towards Multimodal In-Context Learning for Vision & Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:30.950663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:30.950663Z digest=sha256:39a5ba92f84621e20d07be139a1db2d8639ba1e21838b36975e1859e5c097d98

Observation ce4282da-4539-43b6-a236-ed12abc6a68c · inbound

A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges cites this paper.

A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges Towards Multimodal In-Context Learning for Vision & Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:22.904493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:22.904493Z digest=sha256:9743dac609f87d9acf56fa7002593bc46bb406261a4e802d5aa69d9c901a5300

Observation 5d7fb8b5-7c7b-4f08-8ac1-8b13bd9c29eb · inbound

HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains cites this paper.

HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains Towards Multimodal In-Context Learning for Vision & Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:29:52.146055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:29:52.146055Z digest=sha256:e1ebc60574226d1bf34ed6914271cd1080bc4b109a456c01638793d7315c06a1

Observation da1fab29-7ff0-4d7d-800c-e27fe95aa35b · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context Towards Multimodal In-Context Learning for Vision & Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.800157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.800157Z digest=sha256:2f70bb78246e3862447ffa2465af52731642e3434f8922acf0e8980c7c6e6f0c

Observation 7eeb1dd7-260b-4169-8e99-a7e472d65e32 · inbound

Generalizable Object Re-Identification via Visual In-Context Prompting cites this paper.

Generalizable Object Re-Identification via Visual In-Context Prompting Towards Multimodal In-Context Learning for Vision & Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:22.084330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:33:22.084330Z digest=sha256:a0e827e9da86e716ccaca4b2a323558682208785fed08dbc825aad5bc2e2cb10

Observation dcf3cc4d-844a-41ac-a697-cb65eb5830f6 · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Towards Multimodal In-Context Learning for Vision & Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.220406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:274bafd434b1ecb162423a6240de9d8e96ea3579ea30e0bf6e9f975e1c26588d

Observation ea5853cc-6f27-4965-8e07-89a284b14132 · inbound

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models cites this paper.

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models Towards Multimodal In-Context Learning for Vision & Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:24.277663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:24.277663Z digest=sha256:b6efa1c43f0cdc8ee87ef9db0c7eca0affd67c59edd54deef9848a3c1264bbd6