Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:16.518476Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 4 inbound Pith citation observations for arXiv:2506.08967.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:16.518476Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-16T05:59:50.900436Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T05:59:51.133845Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 10cb4c00-fb42-4a17-8643-93c8cb711898 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Claude 3.5 sonnet
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation abd46843-edbf-447c-9d90-54b740e55788 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model ML4CO-KIDA: Knowledge Inheritance in Dataset Aggregation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c04f6727-a3cd-412a-8844-60e880dc20d2 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a5322d-cb5c-40cb-a5c9-1180a7413fd3 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a22e50bd-1de3-49eb-aca5-215897c7b636 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Qwen2-Audio Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0ef7c8-88f3-49f6-a9a9-767451b04f45 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Simple and controllable music generation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 41695169-4cb5-48d0-8092-c55a769324ee · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Recent Advances in Speech Language Models: A Survey
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a60a68-68ec-4beb-b36f-3992a44c7334 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Pengi: An audio language model for audio tasks.Advances in Neural Information Processing Systems, 36:18090–18108, 2023
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04ae4ddb-c8ae-4042-b533-bbee52d4e131 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Kimi-Audio Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 054ed7e2-d761-4403-9623-b840e7a9ba1e · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1daf89f6-b755-43eb-ba41-ef48f608a2c6 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Emo-dpo: Controllable emo- tional speech synthesis through direct preference optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation debb1561-7cd4-450d-94b9-44605811f8d0 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a942c027-9629-45eb-a593-242adb83c897 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Joint audio and speech understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b205fbe5-d36f-4836-bd3b-f5605229ace5 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Gemini 2.0 pro
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f127b7f0-c905-4dec-8d96-fee8ce8bed7d · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d71bed5-e07f-464b-8b25-890a7a7847c6 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa979691-538a-4d86-a543-e88cc561bffa · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Deep residual learning for image recognition
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 82e4a2a7-0850-40ef-9200-49d35d6dd197 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ecf714d-a632-43b9-948e-0244e0216d2f · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ab10ad-e0c6-4cc2-9cb0-80ac60a65f79 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Audiogpt: Understanding and generating speech, music, sound, and talking head
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f417131-a65b-427e-8cfc-91cb7a895586 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GPT-4o System Card
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dfec203-8715-435f-9432-4cfbb285ceb1 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model OpenAI o1 System Card
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681a5cdb-7ba7-4382-9e1f-b936d7d04130 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model WavChat: A Survey of Spoken Dialogue Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b51b3cd-7c60-4af5-b140-ae3258ec7077 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model An llm compiler for parallel function calling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4784308c-b4a6-411a-855a-1dc0db1d4d6d · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d3e3944-1c43-49b2-a452-e36b1b2eda44 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa47b73e-e297-4d3f-a577-85bf14e82af4 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 945b7c88-6470-46fa-989f-fd6bb4e79769 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ac23ad-a909-40b9-81fe-b2e8d0973b32 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Using an llm to help with code understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c60b368-3315-473a-bc21-6a6e4f930483 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model The role of paralinguistic cues in social life.ANALYSIS OF MODERN SCIENCE AND INNOVATION, 1(2):215–218, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2905d971-6f5a-4ca6-b4ba-96df5210b1ce · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model A survey on speech large language models.arXiv preprint arXiv:2410.18908, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd1534b0-420e-4502-8a87-676fc3828313 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2b0cb2-a3ab-483b-adf7-e8e47cc6535a · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd32cd2-70b9-49a6-85db-d6918228071c · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model How to debug code with github copilot
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0450f964-3514-4b32-b6d4-1193c5f26a67 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Paralinguistics in speech and language—state-of-the-art and the challenge
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f9684bb-3fd8-44ed-ba66-efc89785ae4c · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c510c4f-08a8-4eb0-a358-fcd62f230dd2 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Stepeval-audio-360
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 95470915-60f4-4c1e-8152-ab198d7116a5 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d875f8-529e-42f4-811b-7d2830f9f202 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Preference alignment improves language model-based tts
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1195625-63a9-4c2a-b0c9-f18331437a73 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Lami: Large language models for multi-modal human-robot interaction
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 907f9a29-8e0a-4a27-99f6-cd12850217ed · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41321313-bb77-41b8-856f-84ad7fcbe144 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Reinforcement Learning for LLM Post-Training: A Survey
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66595b68-8294-47cb-94a7-89faf54611fd · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Model soups: aver- aging weights of multiple fine-tuned models improves accuracy without increasing inference time
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962da2ef-045a-4b6e-b459-fe49ff99f826 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07125a7f-1975-4d32-a803-21d414475f7b · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model When search engine services meet large language models: visions and challenges.IEEE Transactions on Services Computing, 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f1bb81c-d213-49f6-b752-269636b914d0 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Qwen2.5-Omni Technical Report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd288521-e2f9-45f7-b906-647636d89695 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Uniaudio: Towards universal audio generation with large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5eade042-b8c6-4f85-aa5f-faaa4851b64b · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6364a05c-56a0-4e4d-a795-8ba0b4bfed1d · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0bdad5-184f-4cb4-a36d-75ce3e54ac98 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d49440e-e10f-4e6a-9510-fbd1db9a59de · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359d011e-1276-4af3-b3e7-480e3cd1ae5c · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SpeechAlign: Aligning Speech Generation to Human Preferences
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e36a9ec-f727-4c09-a3e8-6fcb7d392395 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Vistorybench: Comprehensive benchmark suite for story visualiza- tion.arXiv preprint arXiv:2505.24862, 2025
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e967e3e-1770-489c-9c03-e9a1b3228968 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Toolqa: A dataset for llm question answering with external tools.Advances in Neural Information Processing Systems, 36:50117–50143, 2023
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a601aaa1-bf4b-4671-b545-3f298fe43b31 · outbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Pre-trained language model based ranking in baidu search
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77ab2a72-895c-410e-9eb3-653860b0be7d · inbound
Step-Audio 2 Technical Report Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac1ddd35-444d-4687-88c5-1ac7ceb4cc4e · inbound
WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86c944ae-9249-46eb-8c34-1251a9201072 · inbound
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab9d91c2-f249-4a97-9649-5b3a42591fff · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.