Pith. sign in

Paper Citation Record · LEDGER

NExT-GPT: Any-to-Any Multimodal LLM

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 80 inbound Pith citation observations for arXiv:2309.05519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.05519 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 80 of 80 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:29.198766Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:19:20.350619Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3df34cfa-faa2-40dd-8ab2-0d43f0cc918a · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.381448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:9e6d02cf0642465a151796239369b0c75f8dfbe7da88728b45a22b386496b648

Observation 0a8a8608-cd7e-4852-9e53-25befcb570f2 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks NExT-GPT: Any-to-Any Multimodal LLM

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.183385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:0f07139ddc3e2fcb380ecb2e4375022053b062a1a7bd84cb3f46c6a1ac336b49

Observation 42510071-7388-47e6-9155-61ae7828ba4b · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey NExT-GPT: Any-to-Any Multimodal LLM

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:22:56.001147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:c136e23e358a8324c3f543f4006438b1e989e50a95fd4c4b77c5e0be6e069ea4

Observation db948bda-a4a1-4bf3-85e3-84dacd27d500 · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model NExT-GPT: Any-to-Any Multimodal LLM

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.288689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:dda7ad76797d51b87d4c4a9d982fa4b27402be6d5919997922ec908feef83954

Observation 45692325-04a1-46af-8fe8-1c6a85e747e7 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.326697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:61c6c2772269e2ac080a6fd2b82c7ce6337f1197a550dc06022e1e95ce59116d

Observation 2f4d4769-68e1-4ea2-84ef-7833b27cc0eb · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites NExT-GPT: Any-to-Any Multimodal LLM

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.261041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:8e74c3d746d064c654355b737365198ab79bcf285cec3d95393e7ac8a4cdb64c

Observation 082d8ffd-e82a-4384-9e7f-0bc07055dbf8 · inbound

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding cites this paper.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.460606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:db8379bbd95fbe6f79305575324fd7267bfbda7563ce2c5564899cf29b51a345

Observation b2512e14-5851-4017-9d56-dc3c580554a2 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:03:33.624348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:4a6fd8f442de582b1690837958b406de250658b27ef7d2ab2040b763b2802374

Observation 20c6e468-7285-4dc5-a580-52853293193b · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.322185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:bb3ad6119c1b4b0f3b2223cd4e16198b846a71049833b7d9b62ccbcb2213032a

Observation 2720e999-b49d-4523-96c1-3ad73b20fae7 · inbound

Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes cites this paper.

Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes NExT-GPT: Any-to-Any Multimodal LLM

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:43:18.972547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T18:42:00.624044Z digest=sha256:ecc2b85f10839e69dba471d7112b73d41ea3e4f455f056bd09bf165c68e61927

Observation 9fa7f369-9dbe-44e8-97c4-4caaecb6141a · inbound

Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need cites this paper.

Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need NExT-GPT: Any-to-Any Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:35:57.581670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:35:57.581670Z digest=sha256:6eec0a4f4b380454f40a11c02ab862d59e98b0d32dc02abb02cf98e5cf5fef72

Observation 5e680e39-d4f6-4653-8b4c-1eaa2fbf3129 · inbound

Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge cites this paper.

Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge NExT-GPT: Any-to-Any Multimodal LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:19.862100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:19.862100Z digest=sha256:1334f5d19fdf01049c34221fb42f64200891eff1fdd596ad3b66c6b6e379ce40

Observation f9e205e2-78c0-430e-98fa-2615f90cab8e · inbound

Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines cites this paper.

Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines NExT-GPT: Any-to-Any Multimodal LLM

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T13:16:38.339903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:16:38.339903Z digest=sha256:a841e4ceb87524da624bb2574a58f31f6360af5a7b8b9cb206dc2c5879dada9f

Observation 407637a6-f197-4a90-a86f-4f21e97140f0 · inbound

VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension cites this paper.

VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension NExT-GPT: Any-to-Any Multimodal LLM

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T12:17:48.460888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:17:48.460888Z digest=sha256:d334839b7903bd90a3c4b413d8cec176b0585d851225f4d9aa763e1126db8b32

Observation b911cdee-9c6e-4315-8381-836d9152cf18 · inbound

DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users cites this paper.

DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users NExT-GPT: Any-to-Any Multimodal LLM

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:50:51.660025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:50:51.660025Z digest=sha256:174e8a52965d7f9bd0e656be5546e74e3f829fbfd2f914872be1ec646cc5ebfa

Observation 956e5285-307a-47c8-ade5-b529a275ba7b · inbound

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos cites this paper.

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos NExT-GPT: Any-to-Any Multimodal LLM

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T05:53:36.372930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:53:36.372930Z digest=sha256:3f4eeb983f96ae5abd8490bd50863c301dbed28d9d2d82d56c7d2949e20240fd

Observation e32a0a63-e6fb-40fb-8578-477185d6dbee · inbound

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models cites this paper.

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T05:25:25.637079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:25:25.637079Z digest=sha256:2c1fa84d1548702a78427d120b7f5fd7012f765ede1e5284a4edaf2470425b05

Observation a051b0d7-b2bb-4400-8e40-870824ed9ee4 · inbound

EventGPT: Event Stream Understanding with Multimodal Large Language Models cites this paper.

EventGPT: Event Stream Understanding with Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.652501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.652501Z digest=sha256:714349595115c730c63482031b57bc09ad89e653069578dd598ba10996314669

Observation 70807277-4dcf-418b-af70-11f113fbdd21 · inbound

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models cites this paper.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.756955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.756955Z digest=sha256:351d395ef67e2086f9963d45245ee8ccadfb34e4fd71d52bd94e39107d31dfed

Observation 8027ce3b-e69d-497f-8d66-f410dd94008f · inbound

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? cites this paper.

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? NExT-GPT: Any-to-Any Multimodal LLM

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T23:19:11.672010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:19:11.672010Z digest=sha256:b301270aead43ad74a8e132dadd30f0d039aff3ef2275d34cb2b073205c201b1

Observation afb78ad8-e84f-47e9-ba30-fc20ec80c3e9 · inbound

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation cites this paper.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.644671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.644671Z digest=sha256:4b179260b7f97f00001fd51437c08d657292d059abcd6f788ac2d6a9b82acc88

Observation ae4aec21-24c1-4aa6-a82b-ac9ff1b200ac · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.538409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.538409Z digest=sha256:fcdbe67472a5c943d6bc0366a7c6b99416ae1ac324f93073a00c6af9c6542a43

Observation 14c73cd1-b625-4d01-a035-82773801d127 · inbound

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond cites this paper.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond NExT-GPT: Any-to-Any Multimodal LLM

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.342701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.342701Z digest=sha256:4e62bf8c53680e5e96b939815e7a2a29687fb985e89e122565a5e4dff145af92

Observation 7c23ac43-06c9-4272-9ef9-7ad45baa992e · inbound

VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features cites this paper.

VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features NExT-GPT: Any-to-Any Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T19:53:27.465202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:53:27.465202Z digest=sha256:4a94d3de1c2c1fef6417450497ea7f8124cfe2e4da3ff1e5833da58fc598ee59

Observation 326ea139-20b5-4057-994c-1c66eaeec16e · inbound

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models cites this paper.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.980072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.980072Z digest=sha256:b8ad7973a8e70678d9c2db92586bb0dc28a30d228efb00650f2abc1c2d60e819

Observation 147f3c0c-fba2-47d6-abdf-32112d9451b2 · inbound

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding cites this paper.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding NExT-GPT: Any-to-Any Multimodal LLM

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.047135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.047135Z digest=sha256:4bda765aebbcb3ed8be636581b04ea1b8c10a96cf0a7beaa6d2e14ea7447574d

Observation 4573a69c-cef9-4e6e-9e57-06640bc20b12 · inbound

Olympus: A Universal Task Router for Computer Vision Tasks cites this paper.

Olympus: A Universal Task Router for Computer Vision Tasks NExT-GPT: Any-to-Any Multimodal LLM

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.586644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.586644Z digest=sha256:03b48e9dc4e603a8e56861ce59018e17b548e670c7f066a42c2c13dc81686902

Observation 7eadc85c-bbfd-498a-859f-52e931d1f151 · inbound

DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis cites this paper.

DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis NExT-GPT: Any-to-Any Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T14:48:10.820761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:48:10.820761Z digest=sha256:b908dc07fb4f363a2b44a7be40c50722f788edd2a6676b4bbfea30d085441b98

Observation 2353b172-40d4-4f99-947f-172de7282061 · inbound

Deploying Foundation Model Powered Agent Services: A Survey cites this paper.

Deploying Foundation Model Powered Agent Services: A Survey NExT-GPT: Any-to-Any Multimodal LLM

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:46.273534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:46.273534Z digest=sha256:12b17c5b7e87dc863c4dcde749f03e46a85702b304a7b89728b5a6e5c25f8c62

Observation 21ee47f4-540d-4001-9cad-59521d5bb080 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? NExT-GPT: Any-to-Any Multimodal LLM

Reference 199

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:18.001431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:18.001431Z digest=sha256:a31d37845814bac682160710580533160d9bffc79df52740d23f5439d2c3d18d

Observation d199968f-e186-4091-b0d7-e8cb24541128 · inbound

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models cites this paper.

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T06:06:07.105699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T06:06:07.105699Z digest=sha256:cec57cf286e290f16784a86531a25dcba4687a2b1ca7727b33731787424145d4

Observation ef9e7d27-876e-44ac-8e72-8338b337f9a1 · inbound

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues cites this paper.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues NExT-GPT: Any-to-Any Multimodal LLM

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.305307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.305307Z digest=sha256:593dd8d22279a819715c0da1a895b367979f5aa6d6a0b0e616da63bceccf58d7

Observation c5fabe9e-3ec3-4dde-bc8f-b63c067b13ff · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.352028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.352028Z digest=sha256:2eeb11f5a7404a1aba0fe8d17e7982276b271e85857609f38aa39622f92ccc0f

Observation 5ca5ecfb-4837-4432-a4be-f5b66178097f · inbound

Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation cites this paper.

Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation NExT-GPT: Any-to-Any Multimodal LLM

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:53.923845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:53.923845Z digest=sha256:b2bcae0a8ad6536545b34a661036de46b2fdf3393885c6dc78fac0e2f6b05180

Observation 6161e72b-e165-4ef7-8af1-1734474952da · inbound

A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following cites this paper.

A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following NExT-GPT: Any-to-Any Multimodal LLM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:08.782454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:08.782454Z digest=sha256:2ec3917b171bb775fb6191bb11d54fc946c3dec52473d5630c2999eee4a76149

Observation 35656cb0-1090-4d42-9119-aff59f59a1b7 · inbound

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions cites this paper.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions NExT-GPT: Any-to-Any Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.379560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.379560Z digest=sha256:bd8abd9226aa0004701c7aa5ee1d28891f94bb456614b5ff8851b81adb13766b

Observation 0f8b76e2-bc52-4ec0-8e55-7f1438181e21 · inbound

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning cites this paper.

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning NExT-GPT: Any-to-Any Multimodal LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:41.262686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:27:41.262686Z digest=sha256:9512ed1b5b0e7daa1cfb8518d9173ed5a61869f29028c15645cd9d7b024ff8b0

Observation 0364fd52-0414-40b9-aea4-fe1fe809a534 · inbound

Towards Advancing Code Generation with Large Language Models: A Research Roadmap cites this paper.

Towards Advancing Code Generation with Large Language Models: A Research Roadmap NExT-GPT: Any-to-Any Multimodal LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T18:24:12.834124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:24:12.834124Z digest=sha256:8b34dd7da537c9997221bb095b9c1c6997d5f7f2a1d667f647e8d1416ee887b6

Observation 2919fe17-42dd-4e1d-a91c-f17cf84ad08b · inbound

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model cites this paper.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model NExT-GPT: Any-to-Any Multimodal LLM

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.497553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.497553Z digest=sha256:973ce553cd70f5f7c31ac12d43c32aafcfda3b1c3af153beffbd80799cf8c6d1

Observation c7d4997b-4672-4ac8-b692-bd9eac526d5f · inbound

Exploring GPT's Ability as a Judge in Music Understanding cites this paper.

Exploring GPT's Ability as a Judge in Music Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T16:23:18.439080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:23:18.439080Z digest=sha256:20876d6b13f44bdad5034e56df358943f5d38ac0d80d0abd3727761fcdd41e8c

Observation c6389706-4086-4fcb-82d0-f74f613c5f4d · inbound

Parameter-Efficient Fine-Tuning for Foundation Models cites this paper.

Parameter-Efficient Fine-Tuning for Foundation Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:38:02.939869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:38:02.939869Z digest=sha256:06ba247582da832147f6c6accbaf2a4c7907758e0b2d89419d24388e6bf09ac2

Observation d322d50f-c322-4a2d-b23c-5cf795ace02f · inbound

Large Models in Dialogue for Active Perception and Anomaly Detection cites this paper.

Large Models in Dialogue for Active Perception and Anomaly Detection NExT-GPT: Any-to-Any Multimodal LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T13:36:46.982427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:36:46.982427Z digest=sha256:e67922a2d4167ff8d7deb7e99dd3a97a6ecb28248f696f2e09d3a29827520768

Observation 472ea8d0-4ee9-4b21-a93e-59363cad9031 · inbound

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding cites this paper.

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T13:32:58.246447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:32:58.246447Z digest=sha256:e217a1feee552ae8c067de3b799005495fc195ab120df98a437c65a4377dce56

Observation 5f43f488-8b82-49a8-898b-b42f95dcb212 · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling NExT-GPT: Any-to-Any Multimodal LLM

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:14:53.004260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:6433cb2cc6efdee81eed5d0c953b67388692f6254befc61a6f0471f61d0ce486

Observation 871c4183-e5c8-4bfa-98bc-e6d6e3c02401 · inbound

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning cites this paper.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.364061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.364061Z digest=sha256:166d4f43007b5832e06d736e97d6ea667b0baea0932500ebd5169402273db723

Observation eec755af-05c5-4871-b90c-f65d67021776 · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.021100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.021100Z digest=sha256:ff302f17127a32ea9f162fc566a320ed275e1702e63157ad6d12feb571a7f030

Observation cb9d029c-c0d8-463e-80cb-d7253de4152e · inbound

UniCoRN: Unified Commented Retrieval Network with LMMs cites this paper.

UniCoRN: Unified Commented Retrieval Network with LMMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.104506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.104506Z digest=sha256:79f67878dd9bd576bf4286e7c2a337724de3b335da7eac79b537328b4d92f7f4

Observation 02be9b8b-2e7d-40e8-9c38-3db5ebe977c9 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.561275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.561275Z digest=sha256:eff527616a99f4f1b442e8cf8e1f7b1ab1b35905cdc0a5dafc0437758b343f69

Observation adc24064-c836-4590-849d-ead80d009c80 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization NExT-GPT: Any-to-Any Multimodal LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.755345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:3b3356621f5fc9736731cbc2d421f9da9005be089bcbb9ecd52e9880789950fa

Observation 92929f70-6113-4020-9115-e3a98e64831e · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries NExT-GPT: Any-to-Any Multimodal LLM

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.284023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:fc7efacd73a80c60bbb8d1c1ae272ba73f84facc7e41476adf04745ebe8dad59

Observation a07c77f3-61e7-4fa2-9595-d80b5e5c4b9d · inbound

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method cites this paper.

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method NExT-GPT: Any-to-Any Multimodal LLM

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:42.109858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:42.109858Z digest=sha256:ce70ea4c435d1ac46d91665546e1988255c1bc588747f95cc6e6d1fac0439549

Observation 489727a3-afcd-4908-9d9c-f9f8f5e0b8df · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:03.563984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:03.563984Z digest=sha256:5c707115d6c102c0d3781e14d358dcff1a098c51faf39709b1175a30dc251ca6

Observation bfb03666-dc73-483e-a948-ce00f2e52775 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.768622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:135136587857a3edcff540921843e3a38f83f1edbb11289cc188cb41841c21bb

Observation 1e3af57e-dd61-4774-a7a4-6d5ab4db2843 · inbound

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots cites this paper.

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots NExT-GPT: Any-to-Any Multimodal LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:32:19.269257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T13:29:34.151546Z digest=sha256:779f79c58ff32c75e32a78fd106c67d6a8b2b22022a2bc6b61fa2ac55101f66b

Observation f9832be2-b9f5-4fdd-aeac-e089fbae9e01 · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion NExT-GPT: Any-to-Any Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:49.990929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:49.990929Z digest=sha256:6ceb8e81d316a34c197a51d607c3d4f9063463b883f15937c34eb95bf7272d54

Observation 727c3c1d-52ce-4d27-bf43-8b46fbc2fb7f · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning NExT-GPT: Any-to-Any Multimodal LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.460651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.460651Z digest=sha256:93326e24ac2d5b57f76786f89e506ef14bfc76bbc5f3594adb1fc9db08f09a88

Observation 461ee106-6d4a-4934-8cdb-427e0f7fba8f · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.351944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.351944Z digest=sha256:552c4072e5e70ec59ca60cba8ee4bc9e7843385a24d3b4568835a5b81b018dcc

Observation 7684d983-6c48-4c85-9bb7-c98a34ce78cb · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems NExT-GPT: Any-to-Any Multimodal LLM

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.135089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:9866951176e6ce2a0a1c5e93ee0dbeb9c48e55b40dd5a0e7d27e54a4e97fc1d6

Observation 48c745e6-6456-4854-beba-54ef90eaad78 · inbound

Multimodal Representation Alignment for Cross-modal Information Retrieval cites this paper.

Multimodal Representation Alignment for Cross-modal Information Retrieval NExT-GPT: Any-to-Any Multimodal LLM

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:32.375241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:32.375241Z digest=sha256:5bd6adbc7728413eb98c40c988a196ab67f0a7557cd2b16db9c4532ced3938b9

Observation d18c754e-cc81-41f7-95a0-bf11ed74dfb9 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.517877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.517877Z digest=sha256:6fdb80adbc587e53b11336aa80deac2a41841855c5060ec4b44f6ea3ca710862

Observation bfead915-d419-4564-9a55-6c6beaa311aa · inbound

DanceChat: Large Language Model-Guided Music-to-Dance Generation cites this paper.

DanceChat: Large Language Model-Guided Music-to-Dance Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:26:42.724376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:26:42.724376Z digest=sha256:030aa8cbfc311b68dfbb287f4ea0d59b664b90b7c374aeb7e4f2544c4e91a34f

Observation 23ba9004-94bb-42f2-8730-5eecc4029b98 · inbound

Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge cites this paper.

Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge NExT-GPT: Any-to-Any Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:09.376377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:59:09.376377Z digest=sha256:1ef2dbe19dd1ce614397f13ab30efc499f8c040cc0748c103d53504c88e7c731

Observation 3dc49e84-f5b8-4392-8dd6-9ffc27b499d3 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.793866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:dc82fde55fea91a405899c1ed33d654f0a3ca824b4d3c08bdc7b2c65477a4b05

Observation 9950c082-bd59-48a2-b4f5-f46c7fa7a7a6 · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.821240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.821240Z digest=sha256:959aff2d5e50c78e493aa8d85b3a2a273200bfae6d859be8ddeb850a4291630d

Observation d056821a-5da2-4023-a2fa-faf90712d6b6 · inbound

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs cites this paper.

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:33.904321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:33.904321Z digest=sha256:92b9c1c4804219f867cb16fcddde381ddfb37f7a85a0cda306b686ffb4833327

Observation 98fc655e-ce05-4da9-9151-2a93f7c08413 · inbound

DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images cites this paper.

DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images NExT-GPT: Any-to-Any Multimodal LLM

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.162850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.162850Z digest=sha256:0b83dc3112db50f8b5aa883a4f4b91b95a068663438a0779c276a1915327dc22

Observation 9d6818eb-e463-4cf5-8a03-e44237d8a785 · inbound

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan cites this paper.

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan NExT-GPT: Any-to-Any Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:51.608775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:51.608775Z digest=sha256:da033fba6840afd0559ab990cb506440ebf41670d344b827eaa4df0e6df48f29

Observation 0d76b3bd-d2d5-474b-bdc9-8f2c483b45cd · inbound

Audio-Guided Visual Editing with Complex Multi-Modal Prompts cites this paper.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NExT-GPT: Any-to-Any Multimodal LLM

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.833708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.833708Z digest=sha256:85d04be5f2a34a2a6ffdcd367e469bdd2bfefad2b367d85b88ad6e87a27c78b4

Observation 85d0e4a6-b2b6-458c-a997-529ec7b49f2e · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos NExT-GPT: Any-to-Any Multimodal LLM

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:37.086887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:37.086887Z digest=sha256:217e406fc8f7ae11c11ad22b65deaba2154137a5aa844f96e1177145b7ef7a88

Observation b51635ce-46ee-4444-b6c1-0933ccef8253 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.343930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.343930Z digest=sha256:af615f8723288298eebd224be892c736251c0114c8a425e7c6b0d51918fbff9d

Observation d8e3f15a-3507-4bad-b7e8-741d58f22007 · inbound

Cross-Modal Backdoors in Multimodal Large Language Models cites this paper.

Cross-Modal Backdoors in Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.597232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:51:11.432424Z digest=sha256:07e935c39df823f7ffff639270b76cd91384ff95c6c1695ab439ba919948ce5b

Observation 5ed9a3ef-3f09-45c9-ac65-4cb9505567fa · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction NExT-GPT: Any-to-Any Multimodal LLM

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.150954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:944c01194a4e894b94d4f7bd5b800e2dc29df7693766381f19e87825978a6113

Observation cce32fc9-da56-4075-a383-346ea8e357f4 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction NExT-GPT: Any-to-Any Multimodal LLM

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.353196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:ffefd36806a31ba0a0d760085312955785e61dcfb98dfee2dda9c03485660a06

Observation 447a55f8-d548-43a5-b154-53f0f8606604 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.747749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:0b91811026353d9d3dc111e1669064d280873b651d9adf51c684dd2d53bd73c1

Observation c54697af-3353-492b-b83d-ca682f3beb8a · inbound

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions cites this paper.

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions NExT-GPT: Any-to-Any Multimodal LLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T07:49:23.168452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:49:23.168452Z digest=sha256:2d33aee10a9ee80acdaec8206fe84074d85b52a3ab3df7bdb7971baa3eccd14a

Observation 984aec55-eda8-4a39-a2e3-559eded843aa · inbound

Laguerre Geometry for Interpreting Large Language Models cites this paper.

Laguerre Geometry for Interpreting Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T10:42:30.055074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:42:30.055074Z digest=sha256:2fc8f789042d4ed71d66f1bedad602a36f9e087134c92808414385d6ca107972

Observation 0f404c44-6f12-46ba-bceb-19994e9b0001 · inbound

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation cites this paper.

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:27.351494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:38:27.351494Z digest=sha256:b759715fbf411c66d8d2172048ec58563d3bffe2819fd38ebe450c2ac7ab8e8a

Observation d955bed7-164c-439e-ba02-5d2da90d9fc5 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens NExT-GPT: Any-to-Any Multimodal LLM

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:55.978209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:55.978209Z digest=sha256:9505fdc5eabc96f9822ef260fe585af382ca89cfe346906207a04ada5af39674

Observation 3c2af9b0-811a-4563-83a0-7961af6d3030 · inbound

GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation cites this paper.

GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation NExT-GPT: Any-to-Any Multimodal LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:47.192966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:47.192966Z digest=sha256:7bf0c620cf2b7be7de585c1df1c288acdbf5e209c6670e7f0821497bcccc8eb6

Observation d04eb948-bd40-4bd9-90c7-92689792384b · inbound

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design cites this paper.

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design NExT-GPT: Any-to-Any Multimodal LLM

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-14T04:15:29.198766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:15:29.198766Z digest=sha256:f01d3fc3e8863b7c996e5bce4440eca829e352f30fdb926cb7a86c0179a26642