Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 46 inbound Pith citation observations for arXiv:2309.05519.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:21:22.364061Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T01:19:20.350619Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3df34cfa-faa2-40dd-8ab2-0d43f0cc918a · inbound
A Survey on Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a8a8608-cd7e-4852-9e53-25befcb570f2 · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks NExT-GPT: Any-to-Any Multimodal LLM
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 42510071-7388-47e6-9155-61ae7828ba4b · inbound
Large Language Models: A Survey NExT-GPT: Any-to-Any Multimodal LLM
Reference 218
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db948bda-a4a1-4bf3-85e3-84dacd27d500 · inbound
3D-VLA: A 3D Vision-Language-Action Generative World Model NExT-GPT: Any-to-Any Multimodal LLM
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45692325-04a1-46af-8fe8-1c6a85e747e7 · inbound
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f4d4769-68e1-4ea2-84ef-7833b27cc0eb · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites NExT-GPT: Any-to-Any Multimodal LLM
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 082d8ffd-e82a-4384-9e7f-0bc07055dbf8 · inbound
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding NExT-GPT: Any-to-Any Multimodal LLM
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b2512e14-5851-4017-9d56-dc3c580554a2 · inbound
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 20c6e468-7285-4dc5-a580-52853293193b · inbound
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2720e999-b49d-4523-96c1-3ad73b20fae7 · inbound
Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes NExT-GPT: Any-to-Any Multimodal LLM
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f43f488-8b82-49a8-898b-b42f95dcb212 · inbound
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling NExT-GPT: Any-to-Any Multimodal LLM
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 871c4183-e5c8-4bfa-98bc-e6d6e3c02401 · inbound
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning NExT-GPT: Any-to-Any Multimodal LLM
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec755af-05c5-4871-b90c-f65d67021776 · inbound
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding NExT-GPT: Any-to-Any Multimodal LLM
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb9d029c-c0d8-463e-80cb-d7253de4152e · inbound
UniCoRN: Unified Commented Retrieval Network with LMMs NExT-GPT: Any-to-Any Multimodal LLM
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02be9b8b-2e7d-40e8-9c38-3db5ebe977c9 · inbound
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation NExT-GPT: Any-to-Any Multimodal LLM
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc24064-c836-4590-849d-ead80d009c80 · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization NExT-GPT: Any-to-Any Multimodal LLM
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92929f70-6113-4020-9115-e3a98e64831e · inbound
Transfer between Modalities with MetaQueries NExT-GPT: Any-to-Any Multimodal LLM
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a07c77f3-61e7-4fa2-9595-d80b5e5c4b9d · inbound
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method NExT-GPT: Any-to-Any Multimodal LLM
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 489727a3-afcd-4908-9d9c-f9f8f5e0b8df · inbound
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models NExT-GPT: Any-to-Any Multimodal LLM
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb03666-dc73-483e-a948-ce00f2e52775 · inbound
MMaDA: Multimodal Large Diffusion Language Models NExT-GPT: Any-to-Any Multimodal LLM
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e3af57e-dd61-4774-a7a4-6d5ab4db2843 · inbound
Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots NExT-GPT: Any-to-Any Multimodal LLM
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f9832be2-b9f5-4fdd-aeac-e089fbae9e01 · inbound
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion NExT-GPT: Any-to-Any Multimodal LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727c3c1d-52ce-4d27-bf43-8b46fbc2fb7f · inbound
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning NExT-GPT: Any-to-Any Multimodal LLM
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 461ee106-6d4a-4934-8cdb-427e0f7fba8f · inbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7684d983-6c48-4c85-9bb7-c98a34ce78cb · inbound
From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems NExT-GPT: Any-to-Any Multimodal LLM
Reference 197
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 48c745e6-6456-4854-beba-54ef90eaad78 · inbound
Multimodal Representation Alignment for Cross-modal Information Retrieval NExT-GPT: Any-to-Any Multimodal LLM
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18c754e-cc81-41f7-95a0-bf11ed74dfb9 · inbound
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfead915-d419-4564-9a55-6c6beaa311aa · inbound
DanceChat: Large Language Model-Guided Music-to-Dance Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ba9004-94bb-42f2-8730-5eecc4029b98 · inbound
Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge NExT-GPT: Any-to-Any Multimodal LLM
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc49e84-f5b8-4392-8dd6-9ffc27b499d3 · inbound
Show-o2: Improved Native Unified Multimodal Models NExT-GPT: Any-to-Any Multimodal LLM
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9950c082-bd59-48a2-b4f5-f46c7fa7a7a6 · inbound
NeoBabel: A Multilingual Open Tower for Visual Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d056821a-5da2-4023-a2fa-faf90712d6b6 · inbound
Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs NExT-GPT: Any-to-Any Multimodal LLM
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98fc655e-ce05-4da9-9151-2a93f7c08413 · inbound
DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images NExT-GPT: Any-to-Any Multimodal LLM
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d6818eb-e463-4cf5-8a03-e44237d8a785 · inbound
Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan NExT-GPT: Any-to-Any Multimodal LLM
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d76b3bd-d2d5-474b-bdc9-8f2c483b45cd · inbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts NExT-GPT: Any-to-Any Multimodal LLM
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d0e4a6-b2b6-458c-a997-529ec7b49f2e · inbound
Effectively obtaining acoustic, visual and textual data from videos NExT-GPT: Any-to-Any Multimodal LLM
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b51635ce-46ee-4444-b6c1-0933ccef8253 · inbound
Testing chatbots on the creation of encoders for audio conditioned image generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e3f15a-3507-4bad-b7e8-741d58f22007 · inbound
Cross-Modal Backdoors in Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5ed9a3ef-3f09-45c9-ac65-4cb9505567fa · inbound
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction NExT-GPT: Any-to-Any Multimodal LLM
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cce32fc9-da56-4075-a383-346ea8e357f4 · inbound
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction NExT-GPT: Any-to-Any Multimodal LLM
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 447a55f8-d548-43a5-b154-53f0f8606604 · inbound
AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c54697af-3353-492b-b83d-ca682f3beb8a · inbound
EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions NExT-GPT: Any-to-Any Multimodal LLM
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 984aec55-eda8-4a39-a2e3-559eded843aa · inbound
Laguerre Geometry for Interpreting Large Language Models NExT-GPT: Any-to-Any Multimodal LLM
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f404c44-6f12-46ba-bceb-19994e9b0001 · inbound
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d955bed7-164c-439e-ba02-5d2da90d9fc5 · inbound
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens NExT-GPT: Any-to-Any Multimodal LLM
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c2af9b0-811a-4563-83a0-7961af6d3030 · inbound
GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation NExT-GPT: Any-to-Any Multimodal LLM
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.