Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T16:50:43.962905Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2502.01046.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T16:50:43.962905Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-30T22:26:14.531002Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:03:13.782299Z
76 of 76 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5fc86a2a-63b2-486e-896e-48f12d1f1a68 · outbound
Emotional Face-to-Speech write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a82a1ca-ec06-4862-b264-10040089832f · outbound
Emotional Face-to-Speech write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c784a1f-2db3-45d0-982b-b09334faca72 · outbound
Emotional Face-to-Speech LRS3-TED: a large-scale dataset for visual speech recognition
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228fd9bc-ab5a-4b86-8cbf-37803986be1f · outbound
Emotional Face-to-Speech SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1ac5b807-2b4e-45ab-ad6b-f84a0dbfcb44 · outbound
Emotional Face-to-Speech D., Ho, J., Tarlow, D., and van den Berg, R
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bba212cb-3264-4084-8f3b-1641ebdffaea · outbound
Emotional Face-to-Speech W., Fidler, S., and Kreis, K
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4001e787-5070-46db-ab23-5ef9159ade57 · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0ef881d3-5e1b-429f-9015-27998dd27e63 · outbound
Emotional Face-to-Speech and Zisserman, A
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f473141-0a36-416a-a612-0b4510cb8fe6 · outbound
Emotional Face-to-Speech V2C: Visual voice cloning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83a53979-b4a1-44f2-8a66-6c3259c198c0 · outbound
Emotional Face-to-Speech S., Nagrani, A., and Zisserman, A
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 88efc1a8-5fdc-4fe4-89f3-006ecdceea6c · outbound
Emotional Face-to-Speech Learning to dub movies via hierarchical prosody models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9d01dd86-d44e-436e-b922-49c544605764 · outbound
Emotional Face-to-Speech StyleDubber : Towards multi-scale style learning for movie dubbing
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9b997226-46ca-46b9-b267-0d56038f4fa9 · outbound
Emotional Face-to-Speech High fidelity neural audio compression
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 90c26510-fbdd-45a1-ac15-96b6d63c93ec · outbound
Emotional Face-to-Speech Arcface: Additive angular margin loss for deep face recognition
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation df752898-ec6a-4c72-8170-b51272e7ac6d · outbound
Emotional Face-to-Speech and Shutov, V
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 201828de-4436-4332-9442-1cd2f16623d7 · outbound
Emotional Face-to-Speech Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 43a1cc26-d1ee-4d6a-aba2-90c7f0720dbe · outbound
Emotional Face-to-Speech Efficient emotional adaptation for audio-driven talking-head generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6ce5e821-1e50-408b-9497-39e726890cc7 · outbound
Emotional Face-to-Speech Improving adversarial energy-based model via diffusion process
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a6996e8a-99a5-4913-8545-3280c0c30699 · outbound
Emotional Face-to-Speech Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0dcbacbf-7564-47ef-85b0-061428d49176 · outbound
Emotional Face-to-Speech EGC: Image generation and classification via a diffusion energy-based model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7fc0aca9-365b-4dad-96ae-1f9d0521f517 · outbound
Emotional Face-to-Speech Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7870b00c-eed2-4471-bd49-c34b75d4b41c · outbound
Emotional Face-to-Speech An investigation of multi-speaker training for wavenet vocoder
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc283b0d-32bc-419e-82b7-98c5ae23e077 · outbound
Emotional Face-to-Speech and Salimans, T
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fe0ceb16-2d99-4d55-bc37-d4d864504a47 · outbound
Emotional Face-to-Speech Denoising diffusion probabilistic models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b596290c-7592-4d5b-ba2c-cdf081b6ec5e · outbound
Emotional Face-to-Speech and Johnson, L
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0eaf4e25-4cfb-41a3-9591-b6a03d4ab89a · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6f77063-e505-4adc-bede-bbd5ddb52ac3 · outbound
Emotional Face-to-Speech Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1567119-44ca-4a97-bd0a-c4ba830e4cfe · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b7a60f2-fa75-42f0-83c8-bb0b68734653 · outbound
Emotional Face-to-Speech Speak, read and prompt: High -fidelity text-to-speech with minimal supervision
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b425ebf1-8a0f-488a-a629-1bac92c2f7ba · outbound
Emotional Face-to-Speech Deep Directed Generative Models with Energy-Based Probability Estimation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f83871-f69a-45dd-93dd-0b227ece1eae · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 02329e41-fe90-41a0-b8e9-deeb3d81768c · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef8844f7-c50f-409d-abfc-fe157fb86835 · outbound
Emotional Face-to-Speech S., and Chung, S
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7badaced-9882-408d-af97-7de45d9d8ef8 · outbound
Emotional Face-to-Speech Hear Your Face: Face-based voice conversion with F0 estimation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e147df56-4a2d-43a1-a161-0e72be10814a · outbound
Emotional Face-to-Speech UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 694e3347-2324-4734-a331-2f27283a99a4 · outbound
Emotional Face-to-Speech A., Han, C., Raghavan, V
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75876484-b02c-44ce-a574-ebd534e4271b · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4eb7e847-dc45-48ec-a169-667fce75cfed · outbound
Emotional Face-to-Speech Towards a simultaneous and granular identity-expression control in personalized face generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b071e055-cb81-4e53-b369-b905175722f6 · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05cb4b77-1554-4915-a91f-04da4147c73a · outbound
Emotional Face-to-Speech and Hutter, F
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8884668d-d95d-4799-9f20-f312e427d6c6 · outbound
Emotional Face-to-Speech Discrete diffusion modeling by estimating the ratios of the data distribution
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 487895ef-ea72-49bc-aca9-17699ef1d044 · outbound
Emotional Face-to-Speech emotion2vec: Self-supervised pre-training for speech emotion representation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95de6b4f-48c9-4fcd-821b-bae05433aa0c · outbound
Emotional Face-to-Speech POSTER++: A simpler and stronger facial expression recognition network
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7282e835-8cd3-4b1b-a943-84abdcd98820 · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83977cf2-68c9-4aee-8aa0-e5e0e741473a · outbound
Emotional Face-to-Speech Concrete score matching: Generalized score matching for discrete data
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e9dd3b9-8940-4494-a023-8fe45fcc6085 · outbound
Emotional Face-to-Speech HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afdb7d01-8e67-4cb2-a693-ad523e366df3 · outbound
Emotional Face-to-Speech Unlocking Guidance for Discrete State-Space Diffusion and Flow Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21379b9e-a8fb-44e7-953f-b299025db48a · outbound
Emotional Face-to-Speech Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fec4e79-0764-4560-826d-869f4fbe7d28 · outbound
Emotional Face-to-Speech Visual form predictions facilitate auditory processing at the n1
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8b23048-440a-428e-8eeb-b845b9442490 · outbound
Emotional Face-to-Speech and Xie, S
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5ee9044b-52e0-4d3c-b73f-0b813e45db40 · outbound
Emotional Face-to-Speech Hearing faces: Target speaker text-to-speech synthesis from a face
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6f466b0-cfd9-4b98-8ff9-fbdf8debc2f3 · outbound
Emotional Face-to-Speech W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f319cb0e-356d-4538-b860-072a1b7ee870 · outbound
Emotional Face-to-Speech A., Bengio, Y., and Courville, A
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d808a201-85eb-4347-9039-17a490fddc9e · outbound
Emotional Face-to-Speech FastSpeech 2: Fast and high-quality end-to-end text to speech
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b467fb4c-8a20-45dc-97aa-341d8420bd84 · outbound
Emotional Face-to-Speech J., Jin, Q., and Guo, B
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e9a56be-cefb-4282-99bd-0850c636f88e · outbound
Emotional Face-to-Speech Facenet: A unified embedding for face recognition and clustering
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a14e57f9-f64d-4bb3-924b-27c0db7805b6 · outbound
Emotional Face-to-Speech NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e4269af-0220-4bbe-936b-d2ea22cbce0b · outbound
Emotional Face-to-Speech Denoising diffusion implicit models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4f43710-7e79-4477-b3fc-93cce2bdb6cb · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 46ef9fa6-b41a-465a-9401-7e7a96969847 · outbound
Emotional Face-to-Speech Attention is all you need in speech separation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40747989-37d4-44d9-9ee2-83348a6edb91 · outbound
Emotional Face-to-Speech Score-based continuous-time discrete diffusion models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 67357277-30e9-4a4e-867e-96a6c819c725 · outbound
Emotional Face-to-Speech and Fostick, L
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47f5461b-95aa-4381-a1bf-dba0f1d2c3ea · outbound
Emotional Face-to-Speech and Hinton, G
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 51fafd75-1f5b-4a6c-bde3-d68fa5123445 · outbound
Emotional Face-to-Speech and Vanathi, P
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47d97212-12bb-42ba-ba0d-5eb6f0572c2b · outbound
Emotional Face-to-Speech N., Kaiser, L., and Polosukhin, I
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 505ab147-c5c9-486f-80ae-edd969806099 · outbound
Emotional Face-to-Speech Generalized end-to-end loss for speaker verification
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7215bd37-6571-45cc-9288-b4aa56e3d42f · outbound
Emotional Face-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f604beb7-c73f-410f-b68d-c28c69d65f01 · outbound
Emotional Face-to-Speech Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7424b5fb-18ff-4f5e-b46e-3ec3804fc604 · outbound
Emotional Face-to-Speech J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 637cb4a9-52c4-4d75-abe9-81f326f42521 · outbound
Emotional Face-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0137943d-8457-47dd-bd86-8f5564f3e4b2 · outbound
Emotional Face-to-Speech DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1310808-0727-4144-83c7-132e9c8ab306 · outbound
Emotional Face-to-Speech FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a275ec1a-d076-4ea0-ba38-02192b2dc34d · outbound
Emotional Face-to-Speech Diffsound: Discrete diffusion model for text-to-sound generation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7af2d74f-5042-49d0-aaad-507e99131309 · outbound
Emotional Face-to-Speech SoundStream : An end-to-end neural audio codec
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1f177877-1668-4d84-9463-adb26d58ab72 · outbound
Emotional Face-to-Speech SpeechTokenizer : Unified speech tokenizer for speech language models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1302bf96-acf4-4fef-90e8-34fb688272a8 · outbound
Emotional Face-to-Speech Srcodec: Split -residual vector quantization for neural speech codec
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 107ff3ea-8160-4e01-853b-32bfeccd21b6 · inbound
Archon: A Unified Multimodal Model for Holistic Digital Human Generation Emotional Face-to-Speech
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc278dfb-24cf-454f-8d8b-6e0842a3be83 · inbound
Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model Emotional Face-to-Speech
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.