Pith. sign in

Paper Citation Record · LEDGER

Kosmos-G: Generating Images in Context with Multimodal Large Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2310.02992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.02992 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:25.520551Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T15:11:32.537911Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f61ac074-7789-4508-9e7b-ae44c757b176 · inbound

OminiControl: Minimal and Universal Control for Diffusion Transformer cites this paper.

OminiControl: Minimal and Universal Control for Diffusion Transformer Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:41.054515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:41.054515Z digest=sha256:006ba4d1480c6cf4d0106e068eda0a0912f87a3e8f06d9d165efd1383bc42697

Observation 06f58be4-bd22-42bb-ab59-9d4c09cca221 · inbound

DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching cites this paper.

DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:56.078382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:07:56.078382Z digest=sha256:0d7012b84089a23216509f25bd7c168f8e5ae09b0c30a6980579ba2a53a62e0d

Observation b19f6622-8eb3-48cf-869a-8f78ede00d03 · inbound

Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment cites this paper.

Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:35:49.307584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:35:49.307584Z digest=sha256:f9b70238e2e3c9970ada381e5e3dab488d7f046886b7dfa093589204638e1908

Observation 0155ab72-192d-4ae9-bf5c-69e738d3b83b · inbound

Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models cites this paper.

Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:14:56.594042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:14:56.594042Z digest=sha256:e1f7b51e58eb93845c84dded5c0192fa084a1de657c011c68a1b40e5e86df44c

Observation 1947cc8e-98a6-4359-8a23-158730bda8e8 · inbound

Maya: An Instruction Finetuned Multilingual Multimodal Model cites this paper.

Maya: An Instruction Finetuned Multilingual Multimodal Model Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T19:10:57.420110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:10:57.420110Z digest=sha256:0aa06ea945aaf6520441422af1c8c3cd6ac0168ae85929adef34caac9aa7dc01

Observation 4b479e08-728f-4663-b945-eddc8005300b · inbound

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner cites this paper.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.798517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.798517Z digest=sha256:0f6a7b6a0b7730b04bd81a62483f3b162480e39b9ca8cc282a333cf207549996

Observation f839cc3d-e3cd-413b-a1e0-f511a2f4cf8b · inbound

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation cites this paper.

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:24:20.073940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:24:20.073940Z digest=sha256:1fdd11458fedfba277d5e750d5f77fe4e34ecc52d041be120b4cb6d7b56e83f4

Observation 1541a9b8-d780-4dc2-8bc9-5a404e5ef5bd · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 227

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.778741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.778741Z digest=sha256:5cbe0fde28da802c223aa042697f5a02ce32f377145985d385cd11dbed50480a

Observation 7ee35b5d-4d11-4b72-8216-e62c5eb15e9c · inbound

Multitwine: Multi-Object Compositing with Text and Layout Control cites this paper.

Multitwine: Multi-Object Compositing with Text and Layout Control Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.125929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.125929Z digest=sha256:8c96ceeb81ae3b83bbce6de3bc9510ecc28a9ad87b5704d31d14dd587b8df2a5

Observation 2d81f8f4-bbf6-47d5-8c78-f11dd0ff8683 · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.564014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.564014Z digest=sha256:a532ee0c6a8a5c8f96c0ef8ed4c4e7792c743e6add7879aaf1710c5cf4600d0e

Observation 99392501-9814-4cfc-be45-2f80ec30d14a · inbound

Personalized Text-to-Image Generation with Auto-Regressive Models cites this paper.

Personalized Text-to-Image Generation with Auto-Regressive Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:25.520551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:25.520551Z digest=sha256:ea07896c484d4b7d681185885885b0c4d87c2076fb096ce8c67ecbaa86799c93

Observation 28c26d8e-e5d8-4d72-a403-fbff0cdd67ed · inbound

Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions cites this paper.

Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T10:05:17.658110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:05:17.658110Z digest=sha256:7e499a2168e6c28194d32ea4945376b5acb5fe94d17c2721aebf716f3744c3ec

Observation 9c4470ed-be9d-4ca7-b7c9-9ce171088125 · inbound

MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills cites this paper.

MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:51:00.490582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:51:00.490582Z digest=sha256:0a801c036a126d9a40e501a5128725d109af56c36eeaef324db0aa4e32fe8231

Observation 92697fc4-b12a-4c82-99dd-ea89b95b9e8a · inbound

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA cites this paper.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.116160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.116160Z digest=sha256:e050b21a9fee8ee63143d6a341ce1b229e76e4e717b001928dcefb425cdec9bb

Observation 3699ee83-f2a6-4220-9757-7396d850f77e · inbound

STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives cites this paper.

STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:25.394183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:01:25.394183Z digest=sha256:ba6efd02ec05dabef9762d432a23df72b092cc63ed67dac21b4171cbebcde7c6

Observation a22c42b8-9d15-4ba5-ad98-b51135cd7c9d · inbound

Behind Maya: Building a Multilingual Vision Language Model cites this paper.

Behind Maya: Building a Multilingual Vision Language Model Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:22.635956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:22.635956Z digest=sha256:3adfe38b3c89de09ddb748a768381e6bd548f6fb0423dee9760a5fa84409eff7

Observation f90bcb09-fe16-46c2-b5b9-42cba39f935c · inbound

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models cites this paper.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.431664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.431664Z digest=sha256:91cf24791f32a5028b9033e4f0325b7e813dabde8272a30766720fbe651fe629

Observation 0b9d012b-7ae5-462b-905a-907bb39c1e1c · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.998277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.998277Z digest=sha256:abd551cffc61e72c16dcb19b5f8cb7552c1691151b7e8626c4263c3632ba7c79

Observation 996c4bfb-196b-4917-96b6-78e9ee86cc08 · inbound

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM cites this paper.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.316876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.316876Z digest=sha256:18587f89f2acea0de679d017cb14fd15bb163946a9660f73b95307a18a6a170f

Observation bcf822f8-01b4-4cdd-a64d-b01db4ae688a · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:03.103132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:03.103132Z digest=sha256:d459d479a9bb7d608e504d77e028cb1ca19d79fe1494fea06f0a1c16e50a8ed8

Observation c3a3fb72-17da-4008-a3fc-acab907f9fec · inbound

Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects cites this paper.

Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:34.203912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:34.203912Z digest=sha256:dd261024e10c0cfbc6d426c307dea5f94e699074a76a66296395ad9148a48fa4

Observation 578a4aeb-8b48-4569-8d07-50c2268e04d2 · inbound

MultiRef: Controllable Image Generation with Multiple Visual References cites this paper.

MultiRef: Controllable Image Generation with Multiple Visual References Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:32:51.941613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:32:51.941613Z digest=sha256:f1394546ae6b5ed1520aaf8d1329f0cfd4dc41502830454e2ca9652fa11d4e5d

Observation 6b75740f-d7fa-4856-9c42-0d80c079ea97 · inbound

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models cites this paper.

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:38.060925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:38.060925Z digest=sha256:650bb75c9d68ade718f02921f425011960578d9ec0b3635e8a2f4ec81d156533

Observation 196a0af4-6044-4fb0-ae30-1a79c158ce70 · inbound

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy cites this paper.

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T18:37:20.497388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:37:20.497388Z digest=sha256:e4b9b15c4ade02570fd1bf0d1d7114ee88084e29756f47caecf650922baa38ef

Observation de5bd812-9285-47a5-8b1f-221e151053eb · inbound

Animalbooth: multimodal feature enhancement for animal subject personalization cites this paper.

Animalbooth: multimodal feature enhancement for animal subject personalization Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:32.540567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T15:08:36.365653Z digest=sha256:286fa914de81d47ebf35a80886450c370adcac7de5e1c2c1a44e5b6b64a5c04e

Observation d680225d-4a78-4d18-b437-6670990faabe · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:27.197738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:27.197738Z digest=sha256:9e7e1a6977f2394647ac630286dc67e1c5f381a7267a4a461a577907e08921c3

Observation 180e9d82-cd15-43cf-ace4-40e5331e917e · inbound

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing cites this paper.

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:14:08.831286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:14:08.831286Z digest=sha256:74aaf1d2bb11e30a6c9efe4e9839279d5a987e2aa81cd1e148c6d006ab966870

Observation f902196e-3227-4786-8612-9571422c4488 · inbound

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence cites this paper.

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T17:19:38.228209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:19:38.228209Z digest=sha256:4ad783c6d448e1fd77084accd294a370e3caf1751743fa55d63dc01a162ef4f6

Observation 7f858af0-8f98-4295-89b3-e569a54e5417 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 192

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.467697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:08a574e391226b72aacc9c1891cc91350ca1ce494a916effb6a0b17ef2bbba9e

Observation 74bb8080-a6af-4291-9980-545ac4c94525 · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.770498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:55b3820acef47b8ed490b3da0ff13342400603ae735f7d88d04c30d531b95d46

Observation c3140dda-cb53-4d50-a7b9-a704c8732e78 · inbound

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling cites this paper.

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:12.415671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:25:12.415671Z digest=sha256:d46fb6c918fd6caa080a86b41af1b2e5afe3620710dd047e727f3d9a05c5ddf3