Pith. sign in

Paper Citation Record · LEDGER

MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2305.04790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.04790 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:04:40.290744Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

65
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 36daf7e4-f826-41e6-b75a-17c9a5908b9f · inbound

Evaluating Object Hallucination in Large Vision-Language Models cites this paper.

Evaluating Object Hallucination in Large Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:44:09.992326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T13:44:09.626361Z digest=sha256:7829bb39632557fcb8ad39d6f86ce0e274f794c465513c4329ee8c752a888245

Observation 1fc21cfc-c9fd-4b51-bde8-4f916e6eab78 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.533899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:da917844b29cfaabcd164d10560866d0f47e450852f0b0f61152caa489fa5d3b

Observation 86dbe26a-9234-4a96-bcfa-090610f47140 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.661294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:7ba3885f39841e92932efa1ac8f3ea40b980cf216abddf485a1c5d262e3ed5f3

Observation d4a385d0-b631-4b8f-a8a9-6f49cd240e8a · inbound

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning cites this paper.

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:34:56.925616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T17:34:56.836034Z digest=sha256:92bda4c992c3be5984c27fd859d61c6e0efd0e9ef8ddf9b15effadce87ad5579

Observation 6d0cf446-ed1d-46b4-bd04-78d0f90a5ff8 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.218129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:e85365b2be4b54a8fb4eeb794fa03adb33fb945bff00ce3333cb701424bf7b1a

Observation 8d63810e-811b-41fe-b39f-90fe5ddeca6c · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 290

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:54.086446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:c2eb4e3ae69b78cc67cd77be99bd9d4843f1cf48f125672ed6c27f0f5ddcc32c

Observation 5d60023b-8956-4dff-90c9-98a8381dbbe8 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.871504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:b59a46a727f9549c63eca69216535be1566e4dadbf4ab9139db03eadb3140162

Observation cf8190b8-135e-4957-980f-dc37a0eb712d · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:13:08.967416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:aa7d8b1ebfc045f265e4162355185462352699aca7825c137273478bbbb3a82b

Observation 8f444185-6efc-40b1-b169-19812663166e · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.645922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:28cc9f9e9b3bd8398c7e03f8f05f07e847a7907d841b415cc641a0212daaaaf0

Observation d08b84ba-c7ed-4b61-934a-b1d3d5ad9029 · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:08:01.377521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:0b83d9f2f15eff9862bd53859413c24b212ed2a9843c5830fecc2c0dfc3d9611

Observation 3f44ad43-9dd2-4b8c-8bd0-78657d63d7b0 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.127729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ede97b645403b1deaa7803baa4de8c2ae7d560f764086fdc00e4fb24ac3f6120

Observation 57dce32c-9cbd-4f79-92b6-bde2d0787588 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.287822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:e997fe763811430d0643c614627b6e6e514853002c84ebedb542587c5dddfb35

Observation 5c10f305-4288-4827-a656-462842bf09e0 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.115905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:be5e6d56b52ca32c2ac5d3f81faed31df6d92d760cf3f786923c00b29358c46f

Observation ad5d363d-ffeb-4beb-ae4f-f7c953c45272 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.719546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:92a953143188d3ba76294b7dc62c4cb16483ae4c0bbfcf94725b65f8783b81f5

Observation 4f0c569c-2947-40d3-b5a5-39c76f571b76 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.388847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:8aef0c99c42d0ccacbb3b1e7013ea08cd24cff9e8bd031dfb67c8d7523fea9f3

Observation c90589bd-17d0-4111-86b3-031d14f5aefe · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:55:50.308479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:d1ee1687bcc94110911b56ccbefe43c41fc931e1ae1d5fb48ff6b5af79edb50a

Observation 4d673733-08c1-4cc5-85e1-c3dc09dff0a5 · inbound

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving cites this paper.

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:35:42.170256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T16:35:24.063578Z digest=sha256:1f972106dc46b934a5ebffa56728e3720760765537b3526e24bb91926b883295

Observation b78d3394-86f5-4115-8bee-1f12f53155fc · inbound

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective cites this paper.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.290744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.290744Z digest=sha256:cee937e67d10a6aeab8e32e0cd65fa902179cce502515e9abc2e6d724b4e8a8e

Observation 088ced41-4caf-4971-9b92-5c631bac88c9 · inbound

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models cites this paper.

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:40:33.039816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T08:38:02.944230Z digest=sha256:7de5e08d81a851aa484dcc2f07995d8fdd656a94b4142ff235a3c9abac3cc2fe

Observation b6714ece-d05d-4107-9b94-f2872aac866a · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.592607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.592607Z digest=sha256:1615ee9fb8c1682c5772b21f3f6de3fa096cf3fe3722edfacaefa293a1267b2a

Observation 2292cf2d-e482-49cd-afe2-1bcd858ada29 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.397197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:fa50d4c1d41b9f0a7fc04dc08e0e695faaf86e8f71125aae68b5a68b212e6819

Observation 45aeee9e-f93b-48e0-8059-9d25601cfaa0 · inbound

ZINA: Multimodal Fine-grained Hallucination Detection and Editing cites this paper.

ZINA: Multimodal Fine-grained Hallucination Detection and Editing MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:12:14.614209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T10:07:18.162776Z digest=sha256:7a0da84a6614d7ae1e1d070cc6497f3daae035c995873030e81c10a8add5ef5b

Observation 400765dc-0cac-40bc-a2aa-1021cce2aab4 · inbound

MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation cites this paper.

MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:55.367724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:55.367724Z digest=sha256:45863bba701f39c1e9a1383472a05db5e5677bfd75eaea26e9a08050b661bbac

Observation dc773c05-af7f-483e-8f26-a00cc3dd7477 · inbound

EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices cites this paper.

EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:56.603817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:56.603817Z digest=sha256:4017c3f2d1e0b640eed04f22ec52651e99091a003e4d17d858eea7103f849ea4

Observation 5a8ca3ef-4922-450f-8dbc-eab82778dc6d · inbound

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models cites this paper.

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:13.222612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:13.222612Z digest=sha256:761d405ccd0cd55c91cf73b89dfd43fb93a3c42865332908663430f6b2f1456b

Observation a17860a5-add1-408d-b48d-f976276139fc · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:41.699694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:41.699694Z digest=sha256:03ce7bf404080f4affec73ae2d1720d94601d98150091571881f4ac8f08a326b

Observation c873d276-6d9c-4465-aca0-e444a94ae04b · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.629056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.629056Z digest=sha256:0a71e13ab0681503801c1cae6860512095d038792eb3e74fba881e7987c80591

Observation 108f2783-adb1-4ad8-bd18-b35c1718fbf9 · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.673021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.673021Z digest=sha256:fafb493d565fd21e17b9f348bbc76c813d7f7275becfdf3162f76c1c2e714079

Observation 8a31ba8f-7f14-441b-b87f-39e522c77927 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 292

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.434747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:8c066676168fff41cda435d29061c88f3b4603995ef24a9243ba2e20c4663f5d

Observation 149b6ae8-df44-46f2-b627-27211285b6a6 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.569658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.569658Z digest=sha256:40e61a615c37adc3f93a5897e7faa85f3c10e8aab519dc6f0f1ad174602e25fb