Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2305.04790.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T15:04:40.290744Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
65
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 36daf7e4-f826-41e6-b75a-17c9a5908b9f · inbound
Evaluating Object Hallucination in Large Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1fc21cfc-c9fd-4b51-bde8-4f916e6eab78 · inbound
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 86dbe26a-9234-4a96-bcfa-090610f47140 · inbound
A Survey on Multimodal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d4a385d0-b631-4b8f-a8a9-6f49cd240e8a · inbound
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d0cf446-ed1d-46b4-bd04-78d0f90a5ff8 · inbound
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8d63810e-811b-41fe-b39f-90fe5ddeca6c · inbound
The Rise and Potential of Large Language Model Based Agents: A Survey MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 290
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d60023b-8956-4dff-90c9-98a8381dbbe8 · inbound
Improved Baselines with Visual Instruction Tuning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf8190b8-135e-4957-980f-dc37a0eb712d · inbound
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8f444185-6efc-40b1-b169-19812663166e · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d08b84ba-c7ed-4b61-934a-b1d3d5ad9029 · inbound
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f44ad43-9dd2-4b8c-8bd0-78657d63d7b0 · inbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 57dce32c-9cbd-4f79-92b6-bde2d0787588 · inbound
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5c10f305-4288-4827-a656-462842bf09e0 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ad5d363d-ffeb-4beb-ae4f-f7c953c45272 · inbound
Hallucination of Multimodal Large Language Models: A Survey MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4f0c569c-2947-40d3-b5a5-39c76f571b76 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c90589bd-17d0-4111-86b3-031d14f5aefe · inbound
AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 255
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4d673733-08c1-4cc5-85e1-c3dc09dff0a5 · inbound
Visual Adversarial Attack on Vision-Language Models for Autonomous Driving MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b78d3394-86f5-4115-8bee-1f12f53155fc · inbound
Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 142
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088ced41-4caf-4971-9b92-5c631bac88c9 · inbound
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b6714ece-d05d-4107-9b94-f2872aac866a · inbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2292cf2d-e482-49cd-afe2-1bcd858ada29 · inbound
From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45aeee9e-f93b-48e0-8059-9d25601cfaa0 · inbound
ZINA: Multimodal Fine-grained Hallucination Detection and Editing MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 400765dc-0cac-40bc-a2aa-1021cce2aab4 · inbound
MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc773c05-af7f-483e-8f26-a00cc3dd7477 · inbound
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a8ca3ef-4922-450f-8dbc-eab82778dc6d · inbound
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17860a5-add1-408d-b48d-f976276139fc · inbound
"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c873d276-6d9c-4465-aca0-e444a94ae04b · inbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 108f2783-adb1-4ad8-bd18-b35c1718fbf9 · inbound
Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a31ba8f-7f14-441b-b87f-39e522c77927 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 292
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 149b6ae8-df44-46f2-b627-27211285b6a6 · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.