Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:02:51.447632Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 11 inbound Pith citation observations for arXiv:2502.09940.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:02:51.447632Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:53.249841Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
31 of 31 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 3d7b9238-d897-4944-9000-d0126d5d7ae5 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Qwen2-Audio Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c4f458-20b9-42e6-a43b-2c1309aa8a79 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c30636a-ffaf-49f1-8a65-cd9052182336 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Moshi: a speech-text foundation model for real-time dialogue
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34da2de4-288e-48f0-9fb2-dc35f5362924 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd63aac-4a2c-415c-a107-d79b76c623d7 · outbound
A Preliminary Exploration with GPT-4o Voice Mode The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54a85f0-f0cb-403c-84be-bf3409720cca · outbound
A Preliminary Exploration with GPT-4o Voice Mode LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b680d517-637c-45a4-abb2-09fcb9126e8c · outbound
A Preliminary Exploration with GPT-4o Voice Mode WavLLM: Towards Robust and Adaptive Speech Large Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation effb8c2d-a8ee-4f4b-ad31-ee6a915cf8df · outbound
A Preliminary Exploration with GPT-4o Voice Mode Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bdbe4d1-afcd-4684-ba4e-9453112f679c · outbound
A Preliminary Exploration with GPT-4o Voice Mode Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a183cfae-44ac-4211-a996-8cb3c8c1d6ec · outbound
A Preliminary Exploration with GPT-4o Voice Mode Audiogpt: Understanding and generating speech, music, sound, and talking head
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3dc2f48-93e4-4aa5-8bfd-0802688e9f02 · outbound
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1466721-f129-4e76-8e47-4798cfe722b4 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a0e53dda-3bef-4995-a635-8299557f7578 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 34d53109-097c-4089-a55a-918649780fc9 · outbound
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 459e9243-efca-474f-a464-1109ebd917f7 · outbound
A Preliminary Exploration with GPT-4o Voice Mode The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9004b082-8752-4404-b008-c3399c7284ce · outbound
A Preliminary Exploration with GPT-4o Voice Mode Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a253a0-e7b3-4268-8b80-ebdd7e881cf8 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Music understand- ing llama: Advancing text-to-music generation with question answering and captioning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82df58b9-b994-42e6-82d3-15658f617028 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Generative spoken dialogue language modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0982cda7-00bd-4a80-8fda-55c2e3ca948a · outbound
A Preliminary Exploration with GPT-4o Voice Mode GPT-4o System Card
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 756a0851-4d08-4190-87d0-abef841af903 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Training language models to follow instructions with human feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2dadb6-6cfd-4124-b8f1-492cf28c22bd · outbound
A Preliminary Exploration with GPT-4o Voice Mode MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe556de-7c3b-4923-b474-696c96ee8271 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Salmonn: Towards generic hearing abilities for large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c8af6ec6-5620-4688-836b-3fb8d915d8ea · outbound
A Preliminary Exploration with GPT-4o Voice Mode Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ae2e73-8b63-42ed-8ed6-697fe324dfca · outbound
A Preliminary Exploration with GPT-4o Voice Mode LLaMA: Open and Efficient Foundation Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 268babd2-4d67-45e0-8765-793c569cb80c · outbound
A Preliminary Exploration with GPT-4o Voice Mode Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3bfebef-9fda-4a34-99d3-064e69857a8e · outbound
A Preliminary Exploration with GPT-4o Voice Mode Hear: Holistic evaluation of audio representations
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 76c2b940-f58a-4690-9095-1e6e74d0fdd6 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Lin, Andy T
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f7532d53-d73b-4e03-81d0-3dd53e235e36 · outbound
A Preliminary Exploration with GPT-4o Voice Mode Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df5261ba-9ff7-49f2-87df-6dfbafda82cc · outbound
A Preliminary Exploration with GPT-4o Voice Mode Marble: Music audio representation benchmark for universal evaluation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0d2d38f4-6f1f-45b4-a054-2146e245cb87 · outbound
A Preliminary Exploration with GPT-4o Voice Mode SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf9f291e-1e32-4e63-8d09-6d3267b7679b · outbound
A Preliminary Exploration with GPT-4o Voice Mode SpeechAlign: Aligning Speech Generation to Human Preferences
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0222dac1-65b2-4a42-b3f4-ba5a3b496eda · inbound
Breaking the Barriers of Text-Hungry and Audio-Deficient AI A Preliminary Exploration with GPT-4o Voice Mode
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a5ae1d-f3b6-4eb2-8f1b-02d5633e1e7e · inbound
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models A Preliminary Exploration with GPT-4o Voice Mode
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab67a9a-cbc5-4ae2-a27f-821dd93357a5 · inbound
Towards Generalized Source Tracing for Codec-Based Deepfake Speech A Preliminary Exploration with GPT-4o Voice Mode
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e20f3cb-0194-4c2e-9f4e-30aa19a240bd · inbound
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations A Preliminary Exploration with GPT-4o Voice Mode
Reference 208
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ec82f6-b36f-4b8a-a48c-6ad95da979b8 · inbound
When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models A Preliminary Exploration with GPT-4o Voice Mode
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa56d9ba-6427-47f9-997d-c19382bb01b9 · inbound
Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models A Preliminary Exploration with GPT-4o Voice Mode
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b4d92f47-b149-4acb-bda9-9f8c2e7888d7 · inbound
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models A Preliminary Exploration with GPT-4o Voice Mode
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2a7258c4-95a8-4b3f-9e9f-a4f17429687a · inbound
All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation A Preliminary Exploration with GPT-4o Voice Mode
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f148614-e12b-440d-8fc7-495f8f1f40f2 · inbound
ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech A Preliminary Exploration with GPT-4o Voice Mode
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 21f6f1e5-7025-4737-8c84-c0a4aed49abd · inbound
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs A Preliminary Exploration with GPT-4o Voice Mode
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b40423e5-fe69-42c1-84b8-3cdb4d8f8380 · inbound
Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models A Preliminary Exploration with GPT-4o Voice Mode
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.