Pith. sign in

Paper Citation Record · LEDGER

Kosmos-G: Generating Images in Context with Multimodal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2310.02992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.02992 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:24:41.564014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T15:11:32.537911Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2d81f8f4-bbf6-47d5-8c78-f11dd0ff8683 · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.564014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.564014Z digest=sha256:0cbd1a33158547b157aadebcadc62a74270a512f13e08ffbabe68fc64fd93359

Observation f90bcb09-fe16-46c2-b5b9-42cba39f935c · inbound

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models cites this paper.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.431664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.431664Z digest=sha256:376c20b48ec1bbe36cea51b33f71bb80051386b89efad48a9003d61fdbe283c2

Observation 0b9d012b-7ae5-462b-905a-907bb39c1e1c · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.998277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.998277Z digest=sha256:7dae42820bf09cd031957108654e03f8a015332564f6028de9fc29709c8cd3ba

Observation 996c4bfb-196b-4917-96b6-78e9ee86cc08 · inbound

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM cites this paper.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.316876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.316876Z digest=sha256:d930fdb8eeef200cd0fb5ad26c1aa3837b772367000fb9fd9a66604335e251f8

Observation bcf822f8-01b4-4cdd-a64d-b01db4ae688a · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:03.103132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:03.103132Z digest=sha256:dfaa82784db63633ad97a5db6c42beeb785a50b0208219d43d4ac2685285871b

Observation c3a3fb72-17da-4008-a3fc-acab907f9fec · inbound

Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects cites this paper.

Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:34.203912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:34.203912Z digest=sha256:87ad144ad850703d32985dbd1701ae753f2ec8e96354b8559551d54a4a0d202f

Observation 578a4aeb-8b48-4569-8d07-50c2268e04d2 · inbound

MultiRef: Controllable Image Generation with Multiple Visual References cites this paper.

MultiRef: Controllable Image Generation with Multiple Visual References Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:32:51.941613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:32:51.941613Z digest=sha256:5c197c2823238b5a78bc580b5b39936d30e0afee2d06826a6703fc0a3acda905

Observation 6b75740f-d7fa-4856-9c42-0d80c079ea97 · inbound

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models cites this paper.

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:38.060925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:38.060925Z digest=sha256:6281b215b9be1aa6db7882ba91c274614205ad82e48b868678eadacee7238914

Observation 196a0af4-6044-4fb0-ae30-1a79c158ce70 · inbound

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy cites this paper.

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T18:37:20.497388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:37:20.497388Z digest=sha256:7b0ca9c81a79eb19f50133c875d3d71074f6b07408cdaae074de00854a5e4e74

Observation de5bd812-9285-47a5-8b1f-221e151053eb · inbound

Animalbooth: multimodal feature enhancement for animal subject personalization cites this paper.

Animalbooth: multimodal feature enhancement for animal subject personalization Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:32.540567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:08:36.365653Z digest=sha256:385d4be70a6f7848034f11e146c9c18e71a4a6cb323df68d33eef3b7b44f720b

Observation d680225d-4a78-4d18-b437-6670990faabe · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:27.197738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:27.197738Z digest=sha256:b113dd38d19169eb1a3c6e5317b0e243c79046053b8fb60c13ad48e07c65594f

Observation 180e9d82-cd15-43cf-ace4-40e5331e917e · inbound

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing cites this paper.

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:14:08.831286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:14:08.831286Z digest=sha256:3dc3109c1387786fd1863d8e5f2ffe3c41712523bd6eee9861ad6d2cf794998d

Observation f902196e-3227-4786-8612-9571422c4488 · inbound

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence cites this paper.

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T17:19:38.228209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:19:38.228209Z digest=sha256:8c1af9dfb2f93aecedd58f83e0ebb93fc4ca68d1b4d3541b5537875c625ea655

Observation 7f858af0-8f98-4295-89b3-e569a54e5417 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 192

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.467697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:ade210077d63db631c93b28488648674975afe996a9fbabc3ca2d17909b35cc7

Observation 74bb8080-a6af-4291-9980-545ac4c94525 · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.770498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:2924ce468a5289b33443cf38d2c3d76d9db20717794f95dde883e21d8c362ae0

Observation c3140dda-cb53-4d50-a7b9-a704c8732e78 · inbound

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling cites this paper.

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:12.415671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:25:12.415671Z digest=sha256:a0ea834f55e2b5e340090f189cd28b861fb5cb59e75fb643498e42c2716ef7bb