Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:34:02.234077Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2508.16188.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:34:02.234077Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3cf0d4d2-2ebe-432c-9288-0463dec14467 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation LRS3-TED: a large-scale dataset for visual speech recognition
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b47d6a-81c6-486f-bf1e-89fc4ff8527e · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5555e9ba-1fed-44de-8d26-972ac3a39424 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ba25ac1-08ec-4c10-b491-99f59b2ecf7e · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Chang, Sungbok Lee, and Shrikanth S
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75770544-4a51-430d-a404-736e0a4f1e12 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a9401075-ec6d-4da4-9746-ba747df083ea · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Cooper, Michael K
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31389eec-c4ee-457b-9404-3ebfc61d891b · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation VGGFace2: A dataset for recognising faces across pose and age
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f91372ad-c0f2-4ef3-9bd6-62f32321d44d · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Large Language Models are Strong Audio-Visual Speech Recognition Learners
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 835e4410-3f65-4d47-bf3a-bba964b3caf9 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ba0adf-2b9a-40a9-b9e7-0599c93a6649 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Qwen2-Audio Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d293e247-636e-4aec-9648-e436efc70f62 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b876059-c3de-432e-9dc2-a5cf68d9b25c · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6acc3141-b617-4c78-8483-ff80c4b596a9 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5caf015c-fbbf-4dcc-b68c-9331fc6c4e2a · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation High Fidelity Neural Audio Compression
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b15896-bed9-4e2a-8865-9c592fdb7958 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 817dc1c0-57e6-4a7d-b923-1c255ea4fc10 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7597679a-f89b-461f-9337-07358291bb03 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae88353-80e5-4496-9b92-98f7dcf8b312 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d20f6562-f71c-4c16-8d5f-bdf6236aaf4f · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a826c117-363f-48f4-a952-f96380c88037 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation LoRA: Low-Rank Adaptation of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3710c718-f85f-4bb9-8e13-35981f786184 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4bf7298-392e-475e-9afa-343fcd91e00b · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 97e80f69-33b7-42eb-b6dc-6a7bf78de943 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c5553f19-6b4a-4f6c-8e28-6bcc66770cb7 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Kimi-Audio Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1b7d1e-8efa-41bc-ab18-74194098f6d6 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Nicolaou, Athanasios Papaioannou, Guoying Zhao, Björn Schuller, Irene Kotsia, and Stefanos Zafeiriou
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ffd84a08-b0cc-49a2-a138-7d58b13a8840 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d09edf4-5188-44eb-922d-c41a2922ab78 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 663ea35a-21fb-4967-b434-f56e628483c7 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4fc79cfb-f55f-461a-9148-47ecfdc03d70 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d759a2-6f1a-4089-8172-5397774acf31 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cb258ebd-d063-4120-b459-a5dddc0c6ee4 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bdb0781-3549-47ad-99f6-0415d20c507a · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eda168d1-c2e7-434b-a1a9-5e0c7175fa91 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Spirit LM: Interleaved Spoken and Written Language Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa51acb1-b330-48da-b312-6b86216bf9bc · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation GPT-4 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7ad7f9-b680-474e-8659-d0284b47f7ea · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Speech Resynthesis from Discrete Disentangled Self-Supervised Representations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b72384d-209a-43ae-9000-9f2ff11018fe · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Gnana Praveen, Patrick Cardinal, and Eric Granger
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation caf100db-4200-489f-8325-cfe907f14d70 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Robust Speech Recognition via Large-Scale Weak Supervision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b21e640d-fa8a-4b7c-a65a-5a19db3dba99 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation You Only Look Once: Unified, Real-Time Object Detection
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e32739f-246c-40f6-a297-46a552020ec6 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Filntisis, Radek Danecek, Victoria F
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5c9913c5-d99d-4286-9b2d-6fb3bf05cf72 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 110c9b3f-1c72-491a-a7ff-363128b915ad · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Savchenko
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 808bfc5e-bf43-4fc8-9399-8e9f2e1dc2de · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3cf4c4-5941-4935-b7f2-d7a544aa5371 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation SSR: Alignment-Aware Modality Connector for Speech Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4860a1a-7ece-4929-872b-b83f8f6dfb78 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e0c795-db41-4b7d-b1a3-b682878165fa · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c0cd412-a729-4de7-97f9-6016df520be0 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd0e853-501a-417b-9ef9-669af933ccb3 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Learning Emotional Representations from Imbalanced Speech Data for Speech Emotion Recognition and Emotional Text-to-Speech
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e6774275-4465-4540-938b-a465f1af10e6 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation On decoder-only architecture for speech-to-text and large language model integration
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff5aa8c-7edf-46c0-b050-e3613060da67 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f5fa781-25c0-46e3-9844-bdff186365df · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aaea03a9-88b0-4faa-9aa9-509c2580afd1 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0838ca54-6fcb-4d9b-85f9-69feb110e7ab · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Connecting Speech Encoder and Large Language Model for ASR
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe02ded-932b-4f26-995d-3162e8dc2afa · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ccc46f-1fbb-4f3a-a6fc-d4040db498ac · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation online" 'onlinestring :=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712385bd-aedf-40f4-a670-0a09eabc81b9 · outbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation write newline
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.