Pith. sign in

Paper Citation Record · LEDGER

NExT-GPT: Any-to-Any Multimodal LLM

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 46 inbound Pith citation observations for arXiv:2309.05519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.05519 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 46 of 46 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:21:22.364061Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:19:20.350619Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3df34cfa-faa2-40dd-8ab2-0d43f0cc918a · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.381448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:16ebb0e388e42df57c7df42df365f39a77c7654ae49346c223e150918571d7a1

Observation 0a8a8608-cd7e-4852-9e53-25befcb570f2 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks NExT-GPT: Any-to-Any Multimodal LLM

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.183385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:688254303a032b5290a1c34f5efd823a258841f41200ae5877afad07b66bfbe6

Observation 42510071-7388-47e6-9155-61ae7828ba4b · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey NExT-GPT: Any-to-Any Multimodal LLM

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:22:56.001147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:802ba6a0ab383205b550c5aaeeefbe0405b9d74cacb7077bfe114494a868a7fe

Observation db948bda-a4a1-4bf3-85e3-84dacd27d500 · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model NExT-GPT: Any-to-Any Multimodal LLM

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.288689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:58faa798fc28b2db508886280d606201c1bfd3911d7f024fe4e911ed538f5bdd

Observation 45692325-04a1-46af-8fe8-1c6a85e747e7 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.326697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:23a9d2ffb4c6818a8b0aa8d13247b5a73eb41d5a3745190e4b7ce855b3da9da8

Observation 2f4d4769-68e1-4ea2-84ef-7833b27cc0eb · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites NExT-GPT: Any-to-Any Multimodal LLM

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.261041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:e1076363d9f08747368083aceaa7747311d2903ee4b1f35e9745fc1aec243147

Observation 082d8ffd-e82a-4384-9e7f-0bc07055dbf8 · inbound

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding cites this paper.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.460606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:04efca14e84853572a62f161f1d1247be61b1fae5afc780462dfbfd677b7b4fc

Observation b2512e14-5851-4017-9d56-dc3c580554a2 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:03:33.624348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:a2ba393ec46519ec2b5a853870a198c1ff969ec327950621dd3007e35fba28d0

Observation 20c6e468-7285-4dc5-a580-52853293193b · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.322185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:d9e5e0f638f54cc7685630145ed11bd98eed25d68c75c4d6633f08265e7a694d

Observation 2720e999-b49d-4523-96c1-3ad73b20fae7 · inbound

Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes cites this paper.

Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes NExT-GPT: Any-to-Any Multimodal LLM

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:43:18.972547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T18:42:00.624044Z digest=sha256:8e1cfee92271e0505a3e3835b63c6d03b7ba96b4f2357a87e94c5c49de3fca1a

Observation 5f43f488-8b82-49a8-898b-b42f95dcb212 · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling NExT-GPT: Any-to-Any Multimodal LLM

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:14:53.004260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:e35913475c2637bf36d5182e0bc72dbd4685d5bbcbf4ccfc014abae03516e279

Observation 871c4183-e5c8-4bfa-98bc-e6d6e3c02401 · inbound

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning cites this paper.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.364061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.364061Z digest=sha256:7d6bb941b1339f0275d33cad6c1065399fe835838c8872b12353d32b6ba33483

Observation eec755af-05c5-4871-b90c-f65d67021776 · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.021100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.021100Z digest=sha256:0cea86b2cdf967b281ce728e9d5f42d64a4f456169904b26579778645d39367d

Observation cb9d029c-c0d8-463e-80cb-d7253de4152e · inbound

UniCoRN: Unified Commented Retrieval Network with LMMs cites this paper.

UniCoRN: Unified Commented Retrieval Network with LMMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.104506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.104506Z digest=sha256:841b68858e7360d465dee4909218f61db79f934e0bb5f317c119826098fd44ee

Observation 02be9b8b-2e7d-40e8-9c38-3db5ebe977c9 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.561275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.561275Z digest=sha256:c142f2fe5dbc056f523aab8e0cd5f477907d66cfc561e4be98d4b36f18e5f2cb

Observation adc24064-c836-4590-849d-ead80d009c80 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization NExT-GPT: Any-to-Any Multimodal LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.755345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:0fc56d0a30a661f4f80cf836148d05c6fe9cebf449205853a310c9ae40b86770

Observation 92929f70-6113-4020-9115-e3a98e64831e · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries NExT-GPT: Any-to-Any Multimodal LLM

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.284023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:bc2e87c30d33ace8b06359ba080480143de7f3028b73da845055660c7b76825b

Observation a07c77f3-61e7-4fa2-9595-d80b5e5c4b9d · inbound

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method cites this paper.

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method NExT-GPT: Any-to-Any Multimodal LLM

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:42.109858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:42.109858Z digest=sha256:e71ce84ad5b706b805df30b951ae70dd616e2011775541189c5a80cb7f84c286

Observation 489727a3-afcd-4908-9d9c-f9f8f5e0b8df · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:03.563984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:03.563984Z digest=sha256:1a1dfd2256d93fa9ae584621bcc6fd60b3733d79ba4e66210c2960c66c588d2d

Observation bfb03666-dc73-483e-a948-ce00f2e52775 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.768622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:b093530ff0d8e4d1a18efa444ac6a614e46e2411b6e2c1852959272ac6f1d6f8

Observation 1e3af57e-dd61-4774-a7a4-6d5ab4db2843 · inbound

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots cites this paper.

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots NExT-GPT: Any-to-Any Multimodal LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:32:19.269257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T13:29:34.151546Z digest=sha256:ef8c297f2d61bce56cb763d88d408108baf92782647d36248c4a76a971418854

Observation f9832be2-b9f5-4fdd-aeac-e089fbae9e01 · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion NExT-GPT: Any-to-Any Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:49.990929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:49.990929Z digest=sha256:7e7dc141daf9b99c51ee1ed7c0366f91d1f375e0e493d1c7b0730f1e99c74b7e

Observation 727c3c1d-52ce-4d27-bf43-8b46fbc2fb7f · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning NExT-GPT: Any-to-Any Multimodal LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.460651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.460651Z digest=sha256:74143c2538d584676b78ef8e87f1442c5319d131beee8fa36582b9c5dd01249d

Observation 461ee106-6d4a-4934-8cdb-427e0f7fba8f · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.351944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.351944Z digest=sha256:98bb051601124a42e1ab97e4d63f914aa77ce18809506d45a9a4a047afd00e5f

Observation 7684d983-6c48-4c85-9bb7-c98a34ce78cb · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems NExT-GPT: Any-to-Any Multimodal LLM

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.135089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:21e3354b03df35e5d260fc385598a8c0e8530ff5fac1d663be81ac25b66353f5

Observation 48c745e6-6456-4854-beba-54ef90eaad78 · inbound

Multimodal Representation Alignment for Cross-modal Information Retrieval cites this paper.

Multimodal Representation Alignment for Cross-modal Information Retrieval NExT-GPT: Any-to-Any Multimodal LLM

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:32.375241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:32.375241Z digest=sha256:9dd348bdbfd26ce2284d79b3913ea10286d69cd8e5b971583513f4b938150f80

Observation d18c754e-cc81-41f7-95a0-bf11ed74dfb9 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.517877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.517877Z digest=sha256:1746b346321c40b18cfe396ea9ba1885d6b4e3b7468172dddf1bae0024ab247a

Observation bfead915-d419-4564-9a55-6c6beaa311aa · inbound

DanceChat: Large Language Model-Guided Music-to-Dance Generation cites this paper.

DanceChat: Large Language Model-Guided Music-to-Dance Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:26:42.724376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:26:42.724376Z digest=sha256:6622f90fff6deb0a1531b1441e7e356b51c76c8870afa656b125687df42c70f9

Observation 23ba9004-94bb-42f2-8730-5eecc4029b98 · inbound

Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge cites this paper.

Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge NExT-GPT: Any-to-Any Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:09.376377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:59:09.376377Z digest=sha256:898c637d2fb8b081c66a50363beebe6e33c272ac4446c52bc92b1f566859a604

Observation 3dc49e84-f5b8-4392-8dd6-9ffc27b499d3 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.793866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:39b2e04f211b747d9bd15080fcc2128dd57e199cdc81aa78e939d368c07e2665

Observation 9950c082-bd59-48a2-b4f5-f46c7fa7a7a6 · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.821240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.821240Z digest=sha256:cdf1f6a7e8dc69f9c9bff6194db23b69be0ef7cfae42483ca52496c032fa690d

Observation d056821a-5da2-4023-a2fa-faf90712d6b6 · inbound

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs cites this paper.

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:33.904321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:33.904321Z digest=sha256:e33f732ab4baa4885b5d15a2a513eb8b7bec117b930f1517c6e11771a3e0c9e6

Observation 98fc655e-ce05-4da9-9151-2a93f7c08413 · inbound

DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images cites this paper.

DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images NExT-GPT: Any-to-Any Multimodal LLM

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.162850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.162850Z digest=sha256:b2872b97b98c75f09a67faf52f8373e56075bd2ee4b5a1f6b84013ccad93aaf5

Observation 9d6818eb-e463-4cf5-8a03-e44237d8a785 · inbound

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan cites this paper.

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan NExT-GPT: Any-to-Any Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:51.608775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:51.608775Z digest=sha256:6dcadff92f96b56a2a1af589e0dcddad83279fedde00fb01fdb5f0781222bfca

Observation 0d76b3bd-d2d5-474b-bdc9-8f2c483b45cd · inbound

Audio-Guided Visual Editing with Complex Multi-Modal Prompts cites this paper.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NExT-GPT: Any-to-Any Multimodal LLM

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.833708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.833708Z digest=sha256:3c2243601d790c0d6af3d5ce1d94411c1b520c9803798a557fd9f855aecb036d

Observation 85d0e4a6-b2b6-458c-a997-529ec7b49f2e · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos NExT-GPT: Any-to-Any Multimodal LLM

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:37.086887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:37.086887Z digest=sha256:b6114a7792ad4cc7d523ef3b95565165b7a9b262d443eae6517db524b94a645d

Observation b51635ce-46ee-4444-b6c1-0933ccef8253 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.343930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.343930Z digest=sha256:bcb2e605731d35fde8c64141299b9a11bdf9f4e811dea0cf46067c20f0fc7873

Observation d8e3f15a-3507-4bad-b7e8-741d58f22007 · inbound

Cross-Modal Backdoors in Multimodal Large Language Models cites this paper.

Cross-Modal Backdoors in Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.597232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:51:11.432424Z digest=sha256:58ac6c8d9f557f54a5a1fa1b30924af99159fba7a536d673f48c6d6911cea900

Observation 5ed9a3ef-3f09-45c9-ac65-4cb9505567fa · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction NExT-GPT: Any-to-Any Multimodal LLM

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.150954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:1de77112d60eca56b6298a0f49e56c197495e4690c4bdc0920dc7b9839dafa02

Observation cce32fc9-da56-4075-a383-346ea8e357f4 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction NExT-GPT: Any-to-Any Multimodal LLM

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.353196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:0d224b2b096e145fe9e206f6d2d4c071779512ab98d80fc64d61c62862237d6c

Observation 447a55f8-d548-43a5-b154-53f0f8606604 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.747749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:ee809832eac5de02f4f09fefac8bd503b521a04ae8f17e2971cd01f5dc2a905b

Observation c54697af-3353-492b-b83d-ca682f3beb8a · inbound

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions cites this paper.

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions NExT-GPT: Any-to-Any Multimodal LLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T07:49:23.168452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:49:23.168452Z digest=sha256:29a938d5b9c52f6242477d15e0d4cc0351274b287213b60acbc92d8cce69d939

Observation 984aec55-eda8-4a39-a2e3-559eded843aa · inbound

Laguerre Geometry for Interpreting Large Language Models cites this paper.

Laguerre Geometry for Interpreting Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T10:42:30.055074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:42:30.055074Z digest=sha256:163b2b44f5e56767e53461fdc8c543e433117eb35b2a412f3dc5a6b71e59a64b

Observation 0f404c44-6f12-46ba-bceb-19994e9b0001 · inbound

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation cites this paper.

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:27.351494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:38:27.351494Z digest=sha256:8bbb73931195435794453a6266630ddca90113f6480e799061961b69f409a30f

Observation d955bed7-164c-439e-ba02-5d2da90d9fc5 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens NExT-GPT: Any-to-Any Multimodal LLM

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:55.978209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:55.978209Z digest=sha256:b2d1bc0e39172f1e16abbb8cc2ae018dec881dbedf86cdfe8d5da33a17e3e5bc

Observation 3c2af9b0-811a-4563-83a0-7961af6d3030 · inbound

GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation cites this paper.

GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation NExT-GPT: Any-to-Any Multimodal LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:47.192966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:47.192966Z digest=sha256:984b525078e9bce2c7e1f1ec0c753211f55607c49e6e2945fc4eba60e17d9bff