Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:37:25.405095Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2411.13314.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:37:25.405095Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dc4ebf00-4150-4aec-87fe-542f67aadc66 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Tacotron: Towards End-to-End Speech Synthesis
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb9b04be-2296-416d-b495-5f8221359ae1 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Fastspeech: Fast, robust and controllable text to speech,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7c8fb9a9-9428-4907-8a89-668d732f99d2 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fe39d2f6-fcc0-44bd-954b-b91685061290 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee475ef-774e-4f1e-b786-40bf56f2d4f4 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 169871b4-4acf-47c2-a6b8-b227e7d0150f · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Prompttts: Controllable text-to-speech with text descriptions,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9a51dbef-f8e1-434d-b4d1-17e8bc538ca2 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Instructtts: Mod- elling expressive tts in discrete latent space with natural language style prompt,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aca669cf-095d-49cd-95a4-46bece36ba6c · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Mm-tts: Multi-modal prompt based style transfer for expres- sive text-to-speech synthesis,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 75601def-1764-431c-8831-9d72931c4282 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Comparison of different impulse response measurement techniques,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1bb2be7d-8c94-4984-b95e-3187ed4fd516 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Environment Aware Text-to-Speech Synthesis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 637f170c-7109-4377-82ee-998ae92e25df · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception V oiceldm: Text-to-speech with environmental context,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 05883d7e-61ef-4aca-80db-5b1cf7666636 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75755e0f-adfb-430e-b53f-29d599cada24 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Multi-source spatial knowledge understanding for immersive visual text-to-speech,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5830bfac-aef0-4a35-acc2-faef4795b83a · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Learning transferable visual models from natural language supervision,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 26eab282-68aa-40dd-bb41-bbb545d206aa · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Allen, M
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 40f962a1-9329-4381-ae5a-14b7e13fb4bc · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception The festival speech synthesis system,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4296abba-49b2-4dc7-9102-fc2e7cfde937 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception An hmm-based speech synthesis system applied to english,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7d176ddb-9b26-4ff6-935c-4a9d5e2baec5 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b0d3b4-99c5-4976-9e38-e90db2844ca6 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29923ea-f3d5-4bc9-9ff8-efa8f60177a6 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Prodiff: Progressive fast diffusion model for high-quality text-to-speech,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 04fff5d5-02ef-4d4a-8020-d429f7929f39 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Clam-tts: Improving neural codec language model for zero-shot text-to-speech,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b67974d1-4296-4cfe-b0c3-1bff148f1c4e · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Audiolm: a language modeling approach to audio generation,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8685cbbd-04a9-4cf4-8f11-372eed8bc125 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6ad148-2fe3-4a92-b82b-b538d5dbd9c2 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Image2reverb: Cross-modal reverb impulse response synthesis,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 569aa7d9-a979-4b90-bf14-a7dfb4fcf717 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Visual acoustic match- ing,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e95df2cb-c3c2-44db-b18b-71e25a8d0d24 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Self-supervised visual acoustic matching,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e4a73fb9-314e-453b-a55e-b1128a46360d · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Meta-stylespeech: Multi- speaker adaptive text-to-speech generation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57329459-da0e-4a55-9d58-18eb0eea6e71 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Attention is all you need,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71fe08eb-bd37-4259-a8b9-16a3fc517408 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception The lj speech dataset,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96254a7e-cff5-49b0-b8f6-af0bdff4fb4c · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 43e85903-b0a8-459b-b8c9-f919c15bc58b · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Mel-cepstral distance measure for objective speech quality assessment,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 46dacfa0-6931-4e79-81dd-31c6803a785f · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception SC-GlowTTS: an Efficient Zero-Shot Multi-Speaker Text-To-Speech Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4883fca5-8629-48f6-ade2-2c9f514ef95d · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception An overview of voice con- version and its challenges: From statistical modeling to deep learning,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation da33eb71-cba3-4e7f-a4dd-3596581e09f8 · outbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception Grad-cam: Visual explanations from deep networks via gradient-based localization,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.