Pith. sign in

Paper Citation Record · LEDGER

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

As of 4 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 100 inbound Pith citation observations for arXiv:2311.07919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.07919 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T18:57:28.666194Z

measured 137 of 137 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 122 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:17.996500Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:27:36.579792Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact21
  • verified fuzzy13
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f82b18c7-0b96-499e-bba2-fcb3970209e0 · outbound

This paper cites Spice: Semantic propositional image caption evaluation.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Spice: Semantic propositional image caption evaluation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.885441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:e60fc1c468f1c6f9e7f4392c4f67af5fe77cf2f41cecfb4ec587533fc84551a9

Observation 0b34eb66-5d5e-4197-b53c-291d3fa9631a · outbound

This paper cites PaLM 2 Technical Report.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models PaLM 2 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.779045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:18cfba4cb69e2b104155c82c35c9e4960632df6d02b2c264fbac0dba8c88b11d

Observation 1a9a4096-7be8-487e-bd11-85d4f50bebe4 · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.714811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:9b27c49f5e31edd03d0cce3104b014db58bd492173ef60fd6c12672319df8316

Observation 686a88c1-2cb5-4cc3-9743-96b07e7caff0 · outbound

This paper cites Qwen Technical Report.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Qwen Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.736051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:0a31592e0f7a9bd7545cba1d2999bcb8d39aa4dd6d2f14bc29d77f8adc00d2e0

Observation 978cba6e-4754-4698-89bd-496af42f7de5 · outbound

This paper cites AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.831322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:334c380d36580d031da30f34fadeffcbc68be813797573c737cd2f2abb3d6f02

Observation 40075b4d-b472-4db9-9b1e-be1236df4437 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.801899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:0006d2b23fa8aad2703c0a9848e2bad86d6a3046fef09e35b37cf6e94282712d

Observation c2e4f870-f674-4337-aaa5-83566ad92f9f · outbound

This paper cites SpeechNet: A Universal Modularized Model for Speech Processing Tasks.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechNet: A Universal Modularized Model for Speech Processing Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.808161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:401ddc41629ec6968ab65d483370b506a4e4028cfb5997f10ea79be700a8d193

Observation 0389ca6e-f17d-4aec-b5e2-8381638cda49 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.817947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:1bacb1edd077ea2a8734ba71a5130dd0159b8d034a1e55e3b02bf68e4fbcc459

Observation 513cee5e-03c7-4c91-98bd-c697f2efda83 · outbound

This paper cites High Fidelity Neural Audio Compression.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models High Fidelity Neural Audio Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:49:52.306785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:b7f80d0d8a59ff229f0b23a2cb9b2c33cf3a2a3b8b66533c75c1b147e3b94199

Observation 527b4f02-96cd-4cc8-b586-ec04dde22436 · outbound

This paper cites Clotho: an audio captioning dataset.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Clotho: an audio captioning dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.835232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:bf5aebedbcdacd87a3b4784f47f0f5404033de08a66f98efdb094fc6f21b02e0

Observation 68ff7646-e1bf-41d7-bcfa-203ecfe6a788 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.720600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:bba71aa30f5fcaede9389f29e94dd37c96d02fb0a074e6038771b7a5fa199717

Observation 6f5b030d-787a-4213-8656-d5a8ac082472 · outbound

This paper cites CLAP: Learning Audio Concepts From Natural Language Supervision.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models CLAP: Learning Audio Concepts From Natural Language Supervision

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.731393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:8ca835f967368d6b7d36d116a938ec59500a537dcd443166be50ac2d98a4ab6a

Observation 8797769d-2f3c-4613-9ead-011e07268843 · outbound

This paper cites Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.838846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:519f506d83f259975078358738ae8c8e27dcbfc3a9e26c3c95569b31a5ac369a

Observation 3ad2844e-c9f5-4714-bff0-1a4b7959d5fd · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.775426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:e8fd56d5eea593582f13510e714810c0c48b61c89ddaee7e55b0abd45cc9b56b

Observation 36888e76-61f8-40ff-8187-3bcc31335b98 · outbound

This paper cites Vocalsound: Adatasetforimprovinghumanvocalsoundsrecognition.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Vocalsound: Adatasetforimprovinghumanvocalsoundsrecognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.845340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:3f4d876489569a528045348309ce2566b5455fe98c0f239cf702a7e919b65148

Observation f7b85189-ba7f-4e93-8573-4e11a561165d · outbound

This paper cites author Zhou, A.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models author Zhou, A

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.701257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:e13dfda3d6709cc28e64fd0a93741ac33135d71eb4427a134ccdfcc45fd24fd6

Observation 2af46fc4-917a-4c63-8f0e-57d6da47ab1e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.783179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:2aecbf7ea9dd79f82f12a89aab71bdab2c0bc74ae0fe68155d64cdb693f13568

Observation fe563662-f13d-4f5b-860d-d8d85889f490 · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.789265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:b47c6646c72fe60e6a0a7e914b9cf16d1f2710b0ffc09b4fbe6414c6e97b2340

Observation 06221470-b5d6-497b-bdc2-9c210a4d8566 · outbound

This paper cites CochlScene: Acquisition of acoustic scene data using crowdsourcing.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models CochlScene: Acquisition of acoustic scene data using crowdsourcing

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.796789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:abca6a7217bed24eae843e4fed48339ee8ff4ffc301a95e4c3749568b889f926

Observation 59b741bb-9023-4ffb-849d-3efec1ae8228 · outbound

This paper cites an unresolved cited work.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:57:28.849814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:35fad157cc0f676b93ae8ddaa55b93c4e87cb9153aff8bfefdab79c4ca085aa2

Observation 85bb758f-2015-4066-8751-0af96520986f · outbound

This paper cites Clotho-aqa: A crowdsourceddatasetforaudioquestionanswering.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Clotho-aqa: A crowdsourceddatasetforaudioquestionanswering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.859905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:ecf766dcc69dd151afaf2df061bdd1690acb6ccfeb51fcdbe5706ae56917f858

Observation 0e5f935e-91f7-42dd-90d9-b091d4afbb92 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.813519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:2a9ccb14711b6a0eaba016f6bac3a53447b919624f4dba3d5627930900097f0a

Observation 8f2a679e-c6bd-4afd-850c-f503f410178f · outbound

This paper cites Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.823612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:d2ab14acfbfd37bd451db32fb962ee440604ade98f38a573adc7b2ea8dea99f8

Observation 86c9db3c-e837-44ff-a9fc-598806601471 · outbound

This paper cites Montreal forced aligner: Trainable text-speech alignment using kaldi.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Montreal forced aligner: Trainable text-speech alignment using kaldi

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.863538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:ce261af3849ad9464981ec4e27475f7396b210f5b83c99863efd2a1298e016e7

Observation 9138ff39-1341-42d4-9330-8c7a9faac660 · outbound

This paper cites DCASE2017 challenge setup: Tasks, datasets and baseline system.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models DCASE2017 challenge setup: Tasks, datasets and baseline system

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.866991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:053fc9b679b3e6ed731e9dd15092016831b71abe13d4f54635180a5613203896

Observation 7d1d4f61-5064-42a0-8e1d-3aa78034ab6a · outbound

This paper cites Librispeech: AnASRcorpusbasedon public domain audio books.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Librispeech: AnASRcorpusbasedon public domain audio books

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.870304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:82074d92151136388c0b7136b07a215fab6622cefebf569e3afb19ebc0f186d5

Observation 24f53383-474b-4b6d-b5df-d15bf63ea882 · outbound

This paper cites Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.873606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:45e9f749e9f789e4a1b3bc3751397a098817b742ce997f2ab9636b12c58a973b

Observation 4578af68-d155-46d8-8b0e-6e1618754984 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.726199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:2c9d998904e4cd64ce454c55483fe2e0e6726cf917f30660c1d22fa25ee3e0a4

Observation 1b0ec48a-8858-49fd-8f34-0c73b749ae2a · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversations.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models MELD: A multimodal multi-party dataset for emotion recognition in conversations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.876807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:14817f0e86e4416f74ff712ea66169eb54d44a7fa2c4b2249f0cccef6cbd5a33

Observation 5fe6b423-28aa-4c97-9865-085fe94460da · outbound

This paper cites Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.880806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:677d59298ab0e15bfdbf75e8fb11800e84a5a189b7e880907fb835bee8471938

Observation fb10e4d5-8f2d-43af-8bde-5bdd7fc47cd6 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:06:45.816188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:feb4b2daf1361c6d6ae1f83a0d20bb1f108491c7e5f5242dbce3faac6435e291

Observation 2bcc33b7-1fb3-4b57-ad7c-82fe00b902f9 · outbound

This paper cites LLaSM: Large Language and Speech Model.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models LLaSM: Large Language and Speech Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.748529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:6f738d9df436b478d05d923e5e8eb9cdd29a89180f15bd2bb862159f5bfe1d50

Observation d1b73c67-50be-45a3-9491-94e911eddeed · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Emu: Generative Pretraining in Multimodality

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:22:11.462164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:b057a64547bc60b9a72d3acb43e7d949ef49406db8cebe3cf58dca97a53d41fb

Observation 0877f768-621a-4a4f-a3aa-e5cdd7c7ac10 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.757812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:2a312718542dd5129cbba139ccc36858a897d2d40e32c24695b4c0d88a06dab6

Observation 492184fe-5f82-4b06-a1a1-3b6b33fe2ea8 · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.762806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:1ae7e7d97f69031e353f645a926de65fbe8a58e22284115455f848cc095cea13

Observation e09f75fc-0890-4b05-b290-52f56213f9f7 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.769454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:74c35b7aa9e53ea9e6bf78530db703f46c5d3f9ca4b45acb65c7099869b65d31

Observation fe3be170-510f-402f-8ce1-f13f025a6ccb · outbound

This paper cites Whisper-large-v2 Qwen-audio 1st-stage LLM init.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Whisper-large-v2 Qwen-audio 1st-stage LLM init

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.827623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:fb1d41b0f474a63764397ee8961d1edaa0a4e5d28858d53ac2af2500c9694a22

Pith citing papers

Observation 1fce556b-0194-4d5f-9565-dc42f0d4f564 · inbound

Qwen2-Audio Technical Report cites this paper.

Qwen2-Audio Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:14:45.371564Z digest=sha256:15379ee57db78b63e90967fc3e4eae38451683c00f179d18b5595819986ee4ea

Observation 23e023fc-61b4-45fd-a4c8-b6b1a7d5011d · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:50:13.957823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:218732b8225af0a360a72b9dcbfb7385c0ac59034067e871bfc3627279357c5a

Observation af400f69-26f3-42aa-8581-4de09686fd5e · inbound

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot cites this paper.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.586781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:a95ebe2bc1a0edbac39ec7846cf2cc5d418edad71e1cb20696235b42e9a98848

Observation 5a11210e-e87e-4a6c-b8c7-d7f1f4be233d · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:53:26.220877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:eb27bd2ec1d54678a529396c9dacd7e777a5838341bc9140c65b88cb2d5127e3

Observation fee4a322-b101-4a6d-8353-d3e47ddd677e · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 199

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.363845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:16208111c7a96163b13156f23b83b6bb6ca934f6b8afcd238b18f13fac3d6830

Observation f7f4e2d5-eef8-4afe-ac6e-be2886e2d488 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:45:08.215895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:267996e6c2bc767f02ee7bc2d99ddb4f228a1a8daec82b7f8a41f6ae627403a1

Observation 069f554c-bd79-4cfd-a636-58f85cf251ce · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:7dbbcf9e3a9b18f462aa6078257542118f37016af86aac189120ef60f2cf7fed

Observation a68178db-90bf-4a1b-a8a9-5ee404ebe3b1 · inbound

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping cites this paper.

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:41:36.647489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T13:38:17.640841Z digest=sha256:2886c8195e8444070c2f9747d519bb9541963e68f0206b90a5e7907eea8f3553

Observation 8b25ced2-3437-4776-af5b-64709d4dc2ab · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.391139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:02a0a1026ae601fe39dfe20ccc2042575705d62285cd0b83028b53c52fe97db0

Observation 2f947a66-9d89-4f63-8508-05c68e9a5cdb · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.434915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:b374af3c79b055ec4a107e14596843d789aab7353137e65134f1118d6b898842

Observation bb4ccade-bec5-4864-aff7-30c54bb5815c · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:02:16.689598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:23eb068ec9948a105cbbc27278ee7c869e2ccc9f331136ed6ced8670da201f5d

Observation a85c5273-a4ad-4bb3-83da-cfd9fe8cb236 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:59:50.980061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:21b17f11a3668c1bf1f66c0b6bdb9f2d0bd088a8750fd62e364cccbaf1c471dd

Observation 5ba525c0-e879-4e1c-ac0c-dfc514a71efe · inbound

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks cites this paper.

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:41:59.607161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T02:41:52.996457Z digest=sha256:75979878becb1bf8d0d2c9cf493a9a0e94e8a0cbaf88c11feb918832844ac6d1

Observation 9459bfa7-63e8-4767-b177-81c377a0642e · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:24:23.487552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:3eaa986558a58064f2c2a0929b29bab18084b019d9c906738c88278e292dae75

Observation dd84d51e-7157-4d6f-a6f6-17f81ad08e06 · inbound

Direct Simultaneous Translation Activation for Large Audio-Language Models cites this paper.

Direct Simultaneous Translation Activation for Large Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:36.817162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:b3284926a648b44ec8d33fba7f28ae661311215843536e2168c522ecd0b6d55c

Observation a15dbc66-93b9-4f32-825e-c029cb124bea · inbound

GraphMend: Code Transformations for Fixing Graph Breaks in PyTorch 2 cites this paper.

GraphMend: Code Transformations for Fixing Graph Breaks in PyTorch 2 Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:06:35.002150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T16:04:29.215393Z digest=sha256:461ee26d5a330612562568bfac4a4f9ea44be7a47262b16dc66aba27593a363e

Observation f87db4fa-bc05-4d2f-8cab-d25c11610b5e · inbound

Qwen3-Omni Technical Report cites this paper.

Qwen3-Omni Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T00:20:37.406351Z digest=sha256:791f71fdb47a642e39595312032c89f906dd2368f7d3ab2bc41d3db6989bef50

Observation e78fa995-c65f-4887-8e7c-cdeffa3ff2e6 · inbound

Investigating Modality Contribution in Audio LLMs for Music cites this paper.

Investigating Modality Contribution in Audio LLMs for Music Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:50:43.452848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T22:46:49.051113Z digest=sha256:49e56ebde47d17d03c3c80a39a84e494a8f39a1277f1fd63563e71a36298836a

Observation 89840c29-ad96-42f3-868b-da8d84e3413b · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.996500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.996500Z digest=sha256:050aceb7f0c67ed0d88d341cb0ae9f361f4ecf2c89050dd939adf2ce8d2ed7e2

Observation b1907409-d87c-48d9-8cde-3c07821ebec5 · inbound

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents cites this paper.

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:13:42.017859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:13:42.017859Z digest=sha256:e72590ccfe8d5ef7e0677eac0ac45dedc628bb22d74a3587ca30b26ab08dc453

Observation a0e35937-98ce-43e6-bdbc-85b965b995a4 · inbound

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering cites this paper.

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:45:24.311935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T22:44:49.949759Z digest=sha256:c69f8c21bc9cacbe77eda3416a232aa4cf377791ed2bd5a9ef0c6c859015f379

Observation 5a4f7985-4e58-499b-94de-6d9b360f85cf · inbound

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages cites this paper.

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:18:57.209538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:15:04.685150Z digest=sha256:e11fb41d787ef25d42d613c82649dab23475307b0e28822f16753b51d3cdbd2f

Observation b2fca029-712f-4cfe-9d69-cc6c5f6b7842 · inbound

Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching cites this paper.

Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:48:53.066189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:48:53.066189Z digest=sha256:ad79bd1d51b386cd57f367358e48bc6df8c1e5b33d16affbfe12123e05d1e5a5

Observation 938d5d20-137b-438b-92bb-7c6fc392fde4 · inbound

MOSS Transcribe Diarize Technical Report cites this paper.

MOSS Transcribe Diarize Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.056919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.056919Z digest=sha256:24eb848cdf745281db74f2b59a2988cbaf9dedb39615463b78c235afaf3a15cb

Observation 82d211a8-aba7-478b-aa7b-03ef17f646a1 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:01:59.150857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:01:59.150857Z digest=sha256:82ebc5113f62ae33abd52d8e86d54a4ea5bd020e148f546fba37eee14b851ba4

Observation b4e00556-f75b-4533-98b7-ea35e4ccfbd6 · inbound

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering cites this paper.

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:07:58.679402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:04:17.935630Z digest=sha256:0f485f5a382fc53d2fb341a0c29c6e3c978e2794fbd1c08bad18250110d92785

Observation 915656ab-cbcd-4b88-b8c6-f189a0119684 · inbound

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer cites this paper.

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:42.307716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:42.307716Z digest=sha256:5ad63e076a765d7e708519379af834b9b2df1afe2f7108861d4ff8dd52040e57

Observation 736bbc75-f1be-4855-a73d-8c2fb4eef428 · inbound

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity cites this paper.

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:50.902443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:50.902443Z digest=sha256:40bb87dc9814a2d2e96dd55ddf72313e2cb2a5c756926d6734c5c6efd8cf95d8

Observation 4eaee3f0-f2e7-4f2a-89ec-0f84c995de76 · inbound

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling cites this paper.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:41:12.014149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:a28aff9812e068b13b0cc1fc200e1c5174a3aa8173d32c743cbec24684923d46

Observation dec17be9-5a17-4f83-abe7-6e3b2f56bacb · inbound

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision cites this paper.

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:42:05.405658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:42:05.405658Z digest=sha256:9d4a2290c7315f3b84c769a414cee44273f0b3d0f2b99c741c2e45b86cdd26b3

Observation c879398c-5410-45c8-b011-e0edb2ffcbaa · inbound

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips cites this paper.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:ab872099acc5a0abd1aece10f122fea8f73046266f67c6ddef0fd42fe8d951b2

Observation 290cb93c-7d3d-4bcb-bfbd-7e3b0930fc6b · inbound

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition cites this paper.

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:22:08.670559Z digest=sha256:eff8805f7b7f35f31dd5755121494a5be63baa12d6045876ea1c821ca3a86e97

Observation e9a965b2-0d0a-4ef0-9396-243c375e80d4 · inbound

Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training cites this paper.

Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:31:50.670807Z digest=sha256:9b6eec4a40a297545aef1e2b246f7eaaaeec87961307155ab82e502b70b88d86

Observation bd9ea5f8-f59b-491e-ad1c-ff72371925c7 · inbound

Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS cites this paper.

Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:08:56.412939Z digest=sha256:823d94236ca400363fdcaadcabbdce518370e5a119e46fff9bb4deb5740b9890

Observation c749619a-078b-4728-8ab2-c3c25b7369e6 · inbound

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models cites this paper.

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:24:27.118694Z digest=sha256:0107f19426f2d9bf661657087a4ac471b96c2c6340abae6696371342c07de377

Observation c3034d23-7de3-4507-864f-79a5b5ad6212 · inbound

SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding cites this paper.

SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:54:52.275011Z digest=sha256:8799af72db12d9bf4f59426f4f60f1e14e98b3d66fa8d44357aa55dee3fb4ad0

Observation b52494af-c2e3-4ba0-853e-437c93d78287 · inbound

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt cites this paper.

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T12:37:17.843365Z digest=sha256:839db20944cb6b0783fcb02b53a336daff0caa4ae2adda1df36ac4d2d93f5128

Observation 8379af5e-6e66-4f3f-ad8f-1aa9fea82c16 · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:c3bdbc3e99eaee0b20847165404d0bfd12276493ced675f93e9771574cbe006c

Observation 7547d7e1-6daa-43f2-8e63-084d53cb4443 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:d93d79caafa9d12e7d3ff8de740e2dedf517afc29f67686130c92e2222344d52

Observation 5b6f4c32-1675-40bd-987f-f73297d3e4fa · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:cc811ea7e76f99e40c69c76dcff6922b2b8553b7981de4e94409909ec5e5f645

Observation 70ea96b3-229c-4098-bf72-f96e84c1f355 · inbound

Qwen3.5-Omni Technical Report cites this paper.

Qwen3.5-Omni Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:11:22.402552Z digest=sha256:bae949f7fa4697d809ac32a0ce2d4a90c08cd119b101a9dcb3d09128b9c4a8a3

Observation d0052591-51b9-4b3f-9b14-905b6495f83f · inbound

TinyMU: A Compact Audio-Language Model for Music Understanding cites this paper.

TinyMU: A Compact Audio-Language Model for Music Understanding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:17:23.740979Z digest=sha256:50d11b96ed55e9543cf60e78af0abe7240e4707e8d084a8178917fd67567c3fc

Observation 95e3d43a-fc71-4bc5-87b6-3b5faa34f5a9 · inbound

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval cites this paper.

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T03:23:40.993175Z digest=sha256:52ae6af21fe9a3199d913d5eb10833efcbea2a3e5f60c0f2746cef6c5dfbcd39

Observation 2edfa591-0bad-47f3-b02a-3ae6c8c974dc · inbound

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models cites this paper.

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:34:54.375266Z digest=sha256:5f53e6f1db8f59029ab7ccfa64a1067d9bfbb8a6ceccd9eac85daade32e9534f

Observation e5ad04c3-579a-43ea-aa50-76833e525b9c · inbound

Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages cites this paper.

Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T00:34:30.978387Z digest=sha256:183ae685515e9ad078b18732157469c355a125a0dc6c93e547433f8b0d71cb4e

Observation fda59e30-64c1-4c59-bba0-294afcb332f5 · inbound

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence cites this paper.

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T23:03:01.356594Z digest=sha256:0d8f724771997ced68fb3423ce2d0eafe657f7ce8d0049fa412fa8aad2deba2a

Observation 2c5cbd56-3b87-4777-9b2b-6e83b530d87e · inbound

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions cites this paper.

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:25:50.524448Z digest=sha256:1d1e980b4eb0c798449bdc16301394d471d11daa78e457fa8c1fdefd19d7d0a8

Observation abfbf2bf-15ea-4e16-bd20-bb0be9f09530 · inbound

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition cites this paper.

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T18:06:47.859998Z digest=sha256:cfb5d2a17e010369e6ad471c65b6350188f784331e594a1134d7f005fbef7a2d

Observation bdaf2b62-393c-44b1-885c-7505b775a31c · inbound

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition cites this paper.

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T07:25:42.870223Z digest=sha256:8aafd725cb8b98f2d2e907eb26a614e79babc730df8a292fad6cadc618b13a8f

Observation 9e68c7f7-ce71-4e29-b9ae-f44abb081e5c · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:adaf7dae448bd473c1dfa3adcdb185dcf111e255da20847fab7a65c0717eeb71

Observation 5729abf1-bc37-4611-9f1e-e94c51ac38bf · inbound

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes cites this paper.

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T01:14:43.744389Z digest=sha256:9476c073cf12a984cc0ef249ebca4d17378dc9b809a63e213daf906d145c26c1

Observation f5cc76f9-0d30-4fda-a649-90c29cfca4d9 · inbound

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration cites this paper.

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T05:15:55.193062Z digest=sha256:4c218ec18f99eff4abaab2b21ecd37f166e9ce7a150d036c3fefa8bad241ec51

Observation 620ffce7-c856-4ca7-9bb0-3c8f10fadcb3 · inbound

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating cites this paper.

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:17:35.034442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T18:16:22.762405Z digest=sha256:e3ae2e8130316ab63674fcb63f8b0c014586a78ca103f7ff68f651e28f11aa9c

Observation cdc99437-c98d-415e-b494-73f7e9b21483 · inbound

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning cites this paper.

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:24:55.762374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T03:20:37.488507Z digest=sha256:bfc7bcdb73bc44f10cff5c88062d54f5fe2d2379455a66f5f9486be796161dc7

Observation df4a1cc0-e0b4-4445-bf82-31cf959c3e95 · inbound

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction cites this paper.

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:58:13.582853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T10:57:40.997128Z digest=sha256:7f78856bf062e24afb18e757cdd4891001d5458f617b7d1db27bd06538527dd4

Observation 1e53ada0-156c-4bd4-97c7-dbbb9a207388 · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:13:11.904848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:0c2437450d6bec0e1f7cdf9e56658e0ae7afb5bfe7ab0c8b677034f23af5b204

Observation 9e8fa346-5ab2-42b9-a605-1bb28993f172 · inbound

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training cites this paper.

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T07:18:07.283584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T07:13:29.041056Z digest=sha256:b78847002f3e7d2ccadd5ddb4a6108019afece01a117c3c2d94c76c8c06d3853

Observation 4f00b568-ccd7-42fa-9273-d5de3b870bea · inbound

AffectVerse: Emotional World Models for Multimodal Affective Computing cites this paper.

AffectVerse: Emotional World Models for Multimodal Affective Computing Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:48:05.771479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T06:46:33.612905Z digest=sha256:6037f4c99d4d5f85ea71a05a955b6f3c33425ac79a88f9491ad603f6e49ef539

Observation eb3712d9-5bb2-4d85-8236-08592f036432 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.032020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:bfe817a523d9ff9c5574ac5e05d23d1fd0c061805c10ad097a5fa235f60643c9

Observation d0601d4d-da6c-4519-8b2d-1c2ceed0618d · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:44:00.638652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T06:43:52.735211Z digest=sha256:97ffa51c8f24fdd497364c87c9ee421682524d33bb3fb2b9870112618754aeb8

Observation 479df41f-2545-4208-89ef-c1e4ca2aa2a3 · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:45:23.736609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T05:44:46.831360Z digest=sha256:18f215362b6e327e2a08a625dc0c982c481153c8f19e5b6c611a7b2e51645637

Observation 7f4ccaf0-1710-4abd-8708-534d46d21743 · inbound

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods cites this paper.

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:40:53.973928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T01:38:39.499793Z digest=sha256:c26e00febf46e69092418726d0d50ed49a19fc337331426ea20b1233046ce432

Observation 5c25e27e-88b1-45a4-872a-19bd9b4dd46c · inbound

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods cites this paper.

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:44:57.896841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:37:57.477175Z digest=sha256:5ea5998e2d2c69a38e50008305423c9276a42a85584ddc5fdfa44378070064ec

Observation 740e0a2d-bb52-478b-b45b-c7f7d49cb0e8 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.933661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:33bc8807a22764fbf4a84fb95719ad8cc6510f016b5329d5e4414a3683acecea

Observation da3ac5f8-e845-4de2-bfd5-97485eec1690 · inbound

Learning When to Think While Listening in Large Audio-Language Models cites this paper.

Learning When to Think While Listening in Large Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:43:50.761327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T18:37:34.409802Z digest=sha256:dbcfaca819190dcaf95053a1465002c49d99e536323f8f3a72e188c6857541e4

Observation 0bc2763b-f789-42ea-9194-ebae95b4c9f7 · inbound

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation cites this paper.

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T11:43:23.632863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T11:40:46.165547Z digest=sha256:43ee70b05211048301443f0a283134fefc223cc30e79a4ea84539428fa32f4e3

Observation 9d928bee-1a71-4611-b74a-b293c28cb711 · inbound

Decoding Strategies for Diffusion-Based ASR: A Systematic Evaluation of Confidence-Based Thresholding cites this paper.

Decoding Strategies for Diffusion-Based ASR: A Systematic Evaluation of Confidence-Based Thresholding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T05:53:09.221791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T05:45:19.762023Z digest=sha256:6e6d766d3e74fffdaa0fa806461a72b1eeedc70389d04736232a8cd0eaea7223

Observation aefe8a3d-a213-4bab-a83d-2b41b781e83d · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-28T20:52:37.927179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:cc862f027d59f7a713e606bfa45a9d46da56616b0a869cd3846d074f68017779

Observation 179085e3-992f-4d2a-ae0a-f132a28e8a89 · inbound

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors cites this paper.

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:32:35.160558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T19:23:19.422735Z digest=sha256:290b760e770061d5b62522e6d1da5a5c7862e10f7791a1575f413df7268f99b4

Observation f4ca3108-7a9c-4749-b2ad-8a07fd214955 · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.088047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:31cc30a799e688e7f4593876049164d6942f5c16adad7bb99db887b8521675a0

Observation 8849e7a8-7d9b-4497-bb3f-0ad7da2baa51 · inbound

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification cites this paper.

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:06:39.874784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T08:34:11.270002Z digest=sha256:58372ada281b0d9cc4919dae3b9789d775e7d49d335dff27d07b0755d85c2c6c

Observation 723b62a9-8d87-4d49-b458-0387ac1675a5 · inbound

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification cites this paper.

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:24:37.761323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T11:24:22.023188Z digest=sha256:10450be5eac217d796d509994e007b2d31652c179285423bb5810c2d93610439

Observation e2bb7787-a114-40f1-8466-be63fe23eef8 · inbound

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention cites this paper.

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:46:46.222824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T06:41:25.592270Z digest=sha256:56d9dd496dace9709e92b9fce69a172af9d62df550e044876a3e2c16571b60aa

Observation eb536c04-8d0a-4304-9e23-2ec0aa406ed2 · inbound

UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning cites this paper.

UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:16:53.506262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T04:22:13.478562Z digest=sha256:13981f470469f57aa1a0069cb3be1153e62fa9e3f2a7153c9047b608cdf0f2df

Observation a5855a8f-e10c-4cb7-bb7c-fab514158adb · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:46:52.393805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:f07b108c5bb8c6ac22df675efa110a601839457a2d20cbac5470fb43f392ceaa

Observation 932b0ab3-6afd-45de-b862-4d56d7a095e1 · inbound

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models cites this paper.

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:20.020338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T21:16:07.007759Z digest=sha256:b91bf27b0d911fcf55fd468d530f21a7b9af0f082c52ee247a512a48bdf977e7

Observation 3176a1fb-2ea5-4502-bd1c-55e08d7d8913 · inbound

Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation cites this paper.

Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:07:08.932270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T23:00:07.144841Z digest=sha256:312d6984d42a895e6939d6f60bfd71818ef91960032c17ec3049ea89501c94ad

Observation 2eae3c81-fd6b-4d33-881c-12a91f9888d2 · inbound

Titans-as-a-Layer: Test-Time Memory for Conversational Speech Emotion Recognition cites this paper.

Titans-as-a-Layer: Test-Time Memory for Conversational Speech Emotion Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:07:26.859912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T18:28:44.652813Z digest=sha256:fac4fc3fbe57eb59cf0a61e0e3a6a4a9eb6c1abaa1285bffdb39117ffd4ba700

Observation 4ba04366-f80d-42f3-b828-93d0309bccac · inbound

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs cites this paper.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.605027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:1321148fd7b5902771236ed311c0ce05e6773bb2260f0b4c47659cc27d8a4d41

Observation c241abff-92c0-4a8e-9fa7-a78e8cf487c2 · inbound

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales cites this paper.

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.830670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T16:33:41.230574Z digest=sha256:961c8ddc5a97cd8424bc70379b932ac0267ff1862a90a77ea0c19ec8c0f6a7a3

Observation e4925622-4525-4f07-9bf9-2d1dc972c69e · inbound

Speech Encoder Fusion for LLM-based Automatic Speech Recognition cites this paper.

Speech Encoder Fusion for LLM-based Automatic Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T07:57:44.685719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T11:32:58.335783Z digest=sha256:c3a6056d29128572990b7aaa6a4c6ae1054fd948b333927c1db55a03f530bfb9

Observation b6ca04b8-b71d-4992-bd17-b1f28879ad73 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:15:05.345453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:fb096b77cc89b2083ef457bc737cd5b5620bcc923b93665cc82b1658e217dd12

Observation 9452a331-52ee-4638-9edb-3d048f2d355a · inbound

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark cites this paper.

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T07:17:43.937486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T12:11:56.629002Z digest=sha256:5672ca77478c45175e465f3920df39527333bea97f569904aa30687810ecba55

Observation f1b06a76-9818-43c9-900a-f37b5ca039e7 · inbound

DeceptionX: From Multimodal Evidence to Explainable Deception Detection cites this paper.

DeceptionX: From Multimodal Evidence to Explainable Deception Detection Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:27:40.118590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:15:49.796647Z digest=sha256:8ffe4791bf6616beb17d020de772f350c39b5d540f24fe3a7b1ee209b6aee348

Observation 9b7135e4-56ed-4edb-95a9-bb132db28181 · inbound

DeceptionX: From Multimodal Evidence to Explainable Deception Detection cites this paper.

DeceptionX: From Multimodal Evidence to Explainable Deception Detection Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T03:03:32.342985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:03:32.342985Z digest=sha256:b1aa85936fdf05c97ad85d35e6810497c5ab6476010aefa15d8d7a696bd48eb2

Observation ac7ff34f-1db1-4410-b224-25bed8213b59 · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:47:18.010086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:fe3d15193949d77bd9824ac8054318471c21fc651bf8b989dbdc3fb96519cfc3

Observation 619b0362-c950-408b-be0b-f8183fec4b63 · inbound

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era cites this paper.

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:19:44.206175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T11:54:08.457573Z digest=sha256:089ea37cee4203dad3c3d9fc497c22af15671a124d567b21fa6f99da847293a2

Observation 0b2d104a-d72a-4026-8a0a-2d0483a7cfd6 · inbound

Uncertainty-based Debiasing and Unlearning for Decontamination cites this paper.

Uncertainty-based Debiasing and Unlearning for Decontamination Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T12:49:52.436103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T05:57:44.576411Z digest=sha256:0df36f3713ee52a10c870be6863d11558e4c808fcced42caa8c68bd4b0ae50a6

Observation 5ea7e87f-1ab4-41c0-9425-0c9ebd3b1236 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 141

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:20:06.521202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:21f30382c84c9ce2213ff438ad211b91883508da06fb377e0b5201bf8836b756

Observation 4fb3ee76-5ee9-4a24-9bed-c477ba7f6356 · inbound

Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models cites this paper.

Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:30:07.197249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T20:08:33.322532Z digest=sha256:18d55fda7ba3f21c8329d61d8e6833c7c7634f177d7407710856eee78f255d59

Observation 4d7d6bd8-907e-4ffd-8ad1-d78f8f3a39b2 · inbound

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? cites this paper.

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:30:08.192339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T20:03:10.858349Z digest=sha256:2daa6ecca054e2335b7bba72b984150791da20b627e0a1fb6521029e5b16202c

Observation 9df4c430-c924-4e36-9570-2ed681fabf5e · inbound

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval cites this paper.

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:39:58.377508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T02:49:35.819155Z digest=sha256:57498b5f385f77cbd16a4715d6597086ea544080b2691aca31c81b49b87256d2

Observation 83072bef-a268-4225-a828-3e1b98b061ed · inbound

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy cites this paper.

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.453211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T05:06:09.216428Z digest=sha256:cfe43133b3ef12d1edb1147c20aaf6bf7b737da30516e1464d299d47d9a2934e

Observation a1f64d49-bfc4-47c2-8d5c-48b413637a29 · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:54:40.667112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T09:35:03.660671Z digest=sha256:e80a7dd553650fbad37da4d0bede1542359f75caf1bfb88547b086de69a924e5

Observation 89de6386-c342-448a-951d-515551432d5f · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T11:11:08.878850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:11:08.878850Z digest=sha256:6eae6fa61d2e9f6614d0b1a9b22d4fa6f252206cd37ce6b9210c05aade2e6140

Observation 0891cb16-4399-4f0e-b866-501bbb99ede7 · inbound

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs cites this paper.

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:21.799793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:26:38.380118Z digest=sha256:2977a601e95e62704aa1cb93bf4a8a0735fb0b2f24152c6e20f2e9e1619615ba

Observation cc259b6a-0406-45d8-868c-c62268c37d7d · inbound

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models cites this paper.

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T11:55:42.994017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T03:31:42.611472Z digest=sha256:d56d78a123c6b615a4d7cb5c3511c3a5d9a39eee16a782335e4efe9d824a8909

Observation b8215101-0e5e-467c-8a2a-491459e6a730 · inbound

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds cites this paper.

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T03:03:59.776103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T03:03:59.776103Z digest=sha256:05c09db492643024bd3d1dcb65d8e32eb110dc62112986da98498d4f7dbf8b9b

Observation 2f60d2a8-ff7e-4bd3-a57e-289062cf1888 · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T19:34:49.358453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:34:49.358453Z digest=sha256:215933836a3436ada85522c5fe447a51cdfb1369265337d98d546d5c94afec68

Observation ad3f1393-442d-4c5a-b34a-89707874ac6f · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:43:33.132312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:43:33.132312Z digest=sha256:b68269cbdd0423b1bed8979a84626399f4864f60bfaaf34d8845128902574e92