Pith. sign in

Paper Citation Record · LEDGER

Generative Multimodal Models are In-Context Learners

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2312.13286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.13286 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:24:46.417390Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ae3d4092-7407-449d-8aa4-9e8916771767 · inbound

CogVLM: Visual Expert for Pretrained Language Models cites this paper.

CogVLM: Visual Expert for Pretrained Language Models Generative Multimodal Models are In-Context Learners

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.530109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:b66e13e0af54ba5b899192da54940680b7a5d7ee1fd653ea7b27219cc59ae4b0

Observation 6825f400-1343-4766-84a6-68c9cac61b92 · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI Generative Multimodal Models are In-Context Learners

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:41.625490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:aa269b00145be445692cfad07dda32db3fdc94072b56ab49af0af0cb1461fda0

Observation b0911ed6-c864-4058-885a-60fdd5a49a52 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Generative Multimodal Models are In-Context Learners

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.258789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:79678325c4ef794357d8204a908659c144873a0352ed9e0a628a1fc27f87128b

Observation a979d1ca-2f30-44fb-abeb-5bd5a1198235 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Generative Multimodal Models are In-Context Learners

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.444102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:5979dd9f23a9f87eeca9b0e102f3bd256c1862a4a670031b606bd10188d6aa77

Observation c634f635-b37b-4346-92e7-85eb92dcf127 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Generative Multimodal Models are In-Context Learners

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.198403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:c54f96ee7d38cbafc49316b8c2a027803cda08ece91a96c2859971acbb7c4793

Observation 9a5d5d9d-7d52-4730-b776-94eaa623f032 · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Generative Multimodal Models are In-Context Learners

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.412895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:e52056673180949ba589cf1c1b9eb0646995e32975e827a5594ddd1c5ec738fb

Observation 61279890-74ee-4548-bfe9-d7e7817849f0 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Generative Multimodal Models are In-Context Learners

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.569913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:fd8e2e85f32b21dcd5c918d42cb8e698c145f0c31a096d571998447b67a32352

Observation e6da27c2-5c8e-4e1b-a7ba-711c96c4e96b · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation Generative Multimodal Models are In-Context Learners

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:03:33.563332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:26ea1cb1a95ead78f2d58ab3798e37324a63b358862edaf3d4121887aa2dac24

Observation 110289f0-5ca2-4927-8b8a-e23636be0980 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Generative Multimodal Models are In-Context Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.417390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.417390Z digest=sha256:c912432a39d34662204b39d81c2633ea251c253fd550e2a342d0c7299e7c49da

Observation 93fee7ad-4953-479f-a040-2b380c6b26f9 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Generative Multimodal Models are In-Context Learners

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.834228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.834228Z digest=sha256:fd2308d687b8ceae9fc17716936b828c8b784d91c51f92e11b764dedaed250bc

Observation ec11eb0b-6777-46b3-b61b-780cc315fee6 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models Generative Multimodal Models are In-Context Learners

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.902495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:b74743fe597acf6d81e53274b146577c2528e6e8dc29ff71ccbd2d067bdc6af6

Observation c2c0e2d7-3117-4f63-9400-0e3e1d0fe880 · inbound

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation cites this paper.

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation Generative Multimodal Models are In-Context Learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:55.005689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:55.005689Z digest=sha256:d53e0897cba28e1a005ac45e447f2e4d9d833b421ced8b0709c1a880488ef899

Observation 6e9ff967-d6ae-49d2-9586-7c1815053cbd · inbound

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools cites this paper.

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools Generative Multimodal Models are In-Context Learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:36.829877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:27:36.829877Z digest=sha256:ba54160f65f28f5f32a0921b6c2bf5047b6768707c9fdb7fd8004ff72c8aeca6

Observation 4da437f8-b144-4448-a08b-69ee1b9052b8 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models Generative Multimodal Models are In-Context Learners

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.916857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:79dc7ae3c661cfe71a9425e257f360e488231ef97fcbde852a5fe7d755689f74

Observation 7e283520-1b01-40ee-84a0-82bf6d97e28f · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets Generative Multimodal Models are In-Context Learners

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:25.979720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:25.979720Z digest=sha256:8ab4b4cf6e30da15deba97243ed9d191f49b851bb72442b5e2cf484c2051213c

Observation 756d0efe-bc9d-4e89-99d9-89e0a93114cc · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Generative Multimodal Models are In-Context Learners

Reference 245

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.521811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:8249fa9c3c51533a83dfd0279e1d1d5884fcbd5fc98732e188cf331528306a52

Observation d96f5082-d825-49ea-9723-2d8cf0e31c10 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Generative Multimodal Models are In-Context Learners

Reference 154

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:c335ad2ad4bc2f2f8a630cd28f59a325bd4ded56446069fea41135c14bc58d72