Pith. sign in

Paper Citation Record · LEDGER

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 41 inbound Pith citation observations for arXiv:2506.21619.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21619 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:22:00.140173Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:50.544402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:19:49.633371Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba6482a5-3d62-479f-b595-93150211c630 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:51.474748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:51.474748Z digest=sha256:839bf66b7df43a722e64eaadf309464e06841aeb4555988d5a4363071e99dc7e

Observation 1fbb62bb-fe07-419b-8184-9c3d8349b7d2 · outbound

This paper cites C.; Vidler, J.; and Roedig, U.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech C.; Vidler, J.; and Roedig, U

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:09.568747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.714747Z digest=sha256:0b80622e15120950cb65ad433f3c26d512fc8fa1c6f749e80400ee81107c32c8

Observation 857060fd-8691-4b3e-a815-b652d61d99d7 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:09.343035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.891560Z digest=sha256:1a2e41dfc6dd3839b54d2e6dfefc7cd0365d4c1611f0733f944088fcd806eb9e

Observation 55ae0086-9ecf-43fc-b66e-0443261bdd61 · outbound

This paper cites o lge, E.; G \.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech o lge, E.; G \

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:09.141449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.984742Z digest=sha256:ae17c69f30c4a52b8312f7db7311e359f3c9bee027cd95a3de52718dde64f4c8

Observation ed11d40d-2d66-4e37-8525-c8ea8db6105e · outbound

This paper cites T.; Rubanova, Y.; Bettencourt, J.; and Duvenaud, D.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech T.; Rubanova, Y.; Bettencourt, J.; and Duvenaud, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:08.907713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.090822Z digest=sha256:417f1195cb676ed3cce04ea496e60f1afda4ea048c4a873225ad85edb5ffe476

Observation 7c44ba08-afc8-4dbf-94ea-b407c5446dd0 · outbound

This paper cites Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.191734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.191734Z digest=sha256:b34e1e46987acd99ad53b69fda501edfda8c39eba9d864aba5b25626aaf6d3bd

Observation 1e47c365-9195-4e19-8118-c14e329e8df8 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.664334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.374989Z digest=sha256:6d42ece120f9ef7abb905a332c0258624a4a70724d365671609c66dc7a15f0f4

Observation c7f49650-edd8-42f0-9b75-ed6c67c51a73 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.491309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.491309Z digest=sha256:13c33d836984a58676794b8b84d5d4d9749024feb9364144dcd70cfeb6a8e7b9

Observation 3389ed9b-60d1-4229-a949-802ace936876 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.443414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.584839Z digest=sha256:be37fe70c39abcd69ec7cb85d016f5315a7093890405b9bb92c1c86377d42acb

Observation ed3fd843-b8d4-447d-8fe2-389ea4370182 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.224525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.694748Z digest=sha256:debd308f99aa597b1f0689750fd7f41160e9f0e3c60b82320341ad88aa29984e

Observation 743b8f20-f8f9-484d-9729-16dcc0f97530 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.901086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.901086Z digest=sha256:92fc2cfd7738744d9ba770a9d54f9a3a0a78b8bf85bb32dca58ccda765676f6f

Observation 2e9b7afa-6e45-4de3-965f-74d7e921dc57 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:07.977427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.044748Z digest=sha256:bd3251cb91821fc0f8fd23bb1dac8615f7d10ea7cf34610b6ff757c3dd1ad548

Observation eef22172-612c-406c-8739-4ca04512428d · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.174876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.174876Z digest=sha256:d9905763293c16515ba348c8aa83c1fff507be5685fb12b2c0fa94e94e083f8a

Observation 109b26b9-bb32-4ca1-a88b-3a458c154038 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.286323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.286323Z digest=sha256:a265b22e3b137c5e5211023b1336d22d202a39c0715d16ad829e0347d8cae655

Observation 248496b5-51e7-4c45-834d-fdca8bf8d666 · outbound

This paper cites A.; and Wang, H.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; and Wang, H

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:07.662270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.404945Z digest=sha256:1ad71418256f06f87e91be3e06cc65c3b858ccc568a7440a40acc02125f549a6

Observation 68dc8c7f-92f3-436d-8588-5b356d0e5ff6 · outbound

This paper cites E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; Liu, Y.; Zhao, S.; and Kanda, N.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; Liu, Y.; Zhao, S.; and Kanda, N

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:07.397033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.514918Z digest=sha256:8238d053755df9c5e788f426e14ce75833c98363f5e3c6c08a4883d4771ead66

Observation ff85a550-8a16-46de-9218-987d9ba51989 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:07.073937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.654824Z digest=sha256:bcab103dcdb3c36883dbe1ab5b9d20e4acb5cff124a179cfe9b6eaedf3c28abd

Observation 3ecb148e-6917-4155-b0bc-9b5bbfe8e867 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.782697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.794860Z digest=sha256:e97c1cbf8ad6c372cce1e26ec25792c360626584493ddf1c805d1e6dc14808fb

Observation 1112caf4-2c01-4233-bab7-13bf90d6c3d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.970330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.970330Z digest=sha256:cad02bfbf9768f3961e8e59d37980a5309e72516520f5fa8e9ddca8dbf55072e

Observation f04ef65d-a581-47f3-8108-7d8574e282b2 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:54.069634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:54.069634Z digest=sha256:5092944c05a2414e4b89df79fc8f5e2fd15245c3fafa40832146bf2aeae8ff21

Observation c8b94280-cdb3-4f31-a909-433f45771f95 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.449439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:54.324773Z digest=sha256:95a29cc57767ed4e99081a35c29d7b0003050c6f6e6c6deaaf60160788273c3c

Observation e314fc3a-0310-4ba0-888d-305ab9815519 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:54.509932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:54.509932Z digest=sha256:c635629f0d0f4c9c764ade8eb4b892d04931c59b58de7fe89995b01effdbb4af

Observation 15ff03d1-ba1d-4f90-b261-66173c4e3946 · outbound

This paper cites ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:22:01.640762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:54.594770Z digest=sha256:3b03031fa217a21563a5c8d2546da55d761601384c239bcea8a8ad2b881348da

Observation a83a6da2-1f0f-4950-84ee-9026ec8c28eb · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.093152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.093152Z digest=sha256:695f0de85efdbf6f0f7e33058884912a95cc8d414fdf04dd3f73b41fbfc3408a

Observation a2e58924-eafc-479b-b71f-e7919cd052b9 · outbound

This paper cites SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.224747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.224747Z digest=sha256:b46e4263949ca93b73cc092e53ebed3bc5d95e9514745e55ee189e919856c402

Observation 50297084-4b08-47df-8773-dffb35428c2b · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.135488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.490696Z digest=sha256:493cd210182f7f95a431c68ae4aba707e9b008cb5a1d5ae8d7ae6fbac08cca8d

Observation eb32930b-3b78-44ad-baa0-fadb9101c5b4 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.968350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.624763Z digest=sha256:bca723af17fd879c8a6ad6a6d85c401422ee40c06e13bccc64050cd163567d6c

Observation 0fe15248-fe58-4e00-8cc3-3c21292f24be · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.719910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.719910Z digest=sha256:ce00ae5b82f2877b805ca2f3ae05d237d3d329186894520af002e224d70d8d59

Observation 76fed29e-242d-4dfa-b6fc-41b2f8fc9140 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.843335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.843335Z digest=sha256:eb3136ead2f4a9e76004bd982c23558f39f4c6b0b593b3c01d3587cb470b5192

Observation 55540a81-a7ae-44cd-91a1-fd5e4d506dc2 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.796837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.959394Z digest=sha256:951fb60d89cb3ec24ed22ed77eba8cb67634e6090d4ecee4e8ef5c47ac6dc4fc

Observation ec264d95-654f-440a-9c15-2b2f25920f8a · outbound

This paper cites FleSpeech: Flexibly Controllable Speech Generation with Various Prompts.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.065145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.065145Z digest=sha256:cc998a8aff78c62c3b48e24d26d7df17adb0d7088f5d12dfff66665b0fe9a66f

Observation 2b651a00-37be-422f-bbfc-7bbf58c16c60 · outbound

This paper cites A.; Han, C.; Raghavan, V.; Mischler, G.; and Mesgarani, N.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; Han, C.; Raghavan, V.; Mischler, G.; and Mesgarani, N

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:05.560496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.158020Z digest=sha256:25cd9128007c1032f04dac1c950e7403566c6e92435338253fcccf2ddd3c1372

Observation 2bdd39a2-b877-43b9-ba76-26e40715669f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.244739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.230762Z digest=sha256:7a8e214c09065bc3c67ad3e1df1125aebc844286a7cef3e60c4352989483580f

Observation d4926be1-b5a6-4c69-9098-2f81b5e375b4 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Zero-shot Voice Conversion with Diffusion Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.314992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.314992Z digest=sha256:9b05c84c3695791a2d1fe310e271aee8774920894c963fb6f855cfc5479aaf2a

Observation db8cd0e8-bcf4-4dc7-9249-62ea0776d5a3 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.820065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.474829Z digest=sha256:4c68b40f5d748dbf29bc0a8cf78334b8ff74152e97c83b95c6fb5777b1c45be2

Observation d7051c0b-f641-4363-940d-13858c01e34d · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Finite Scalar Quantization: VQ-VAE Made Simple

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.557191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.557191Z digest=sha256:fd5893c1529d6d91d9e8766f135ce5725b657c8ec541535a9a8fa6b7c84cfee1

Observation 25004936-e1d2-498e-a7ea-46875215afb1 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.630048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.630048Z digest=sha256:48bf2b9ab882d2d853ba0bdb5adeaceaf86fec65f2452edcf33fdf3b78a859ad

Observation 42046312-291d-4ca5-8cb5-79b7fa1cdd94 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.720288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.720288Z digest=sha256:51797b65af864f9bd4c7e55cb7b81f391328ed90858f3543b48ac9e02728d207

Observation ca11a80b-9add-4a97-88fa-d01d3087f14f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.538604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.085294Z digest=sha256:56f6171fb297f077d7d18110162e375a9f72702b75ae87993aacdb37a9504d6f

Observation 51dd94e3-a1a6-48c6-9188-206b659b9008 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:04.386285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.212866Z digest=sha256:df671d27cd296fbe3a8d9e79d890926ad58b1e3a8f5e497cfb670697f725bccc

Observation 9ecc25f5-0aeb-4093-ba3e-7273c48a84ff · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:57.367656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:57.367656Z digest=sha256:f8bb21f4696cd517c7907ed889687d83f70688e3e4bd73221ca915bf79d0bc7a

Observation 4753148e-6bf7-473f-8d88-4f3c673be2e8 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.191326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.484383Z digest=sha256:e02f5957c0b0e94cb909fbee5a4d7274b1766bfc677c5f95e8ae14e56c7000f5

Observation b6bf2331-1b89-4039-938e-dc3b51fde4bb · outbound

This paper cites A.; Gonzalez, J.; and Escalera, S.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; Gonzalez, J.; and Escalera, S

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:04.028061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.624749Z digest=sha256:a2a51e905cfeda1aacabd7ff63d34a97bacb555ecc28b43261f8c5725353d8ec

Observation f6c00cfd-550f-4e87-9953-321632221f2c · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:57.740164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:57.740164Z digest=sha256:95e5088cdc86978a681d678ef4fe215e12b5ef9cb0ebee5d90ac1bd8c39556b9

Observation d58f3e58-36fc-4fd0-8dcd-579b40551c1b · outbound

This paper cites E.; Hinton, G.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech E.; Hinton, G

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:03.851801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.864744Z digest=sha256:9a3957d0af2db79f5df4c7ec7c3001739552039b07d5f6731533077251f52813

Observation 214db0ad-301f-4718-9fc1-d18dda973ad7 · outbound

This paper cites DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:22:01.078010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.044747Z digest=sha256:d85805af7f7b0b3f52d3a530881fe909b14929af870df4a11c0055eae90f0ee9

Observation e3098054-309b-4bc4-bf24-f287f3b7d03d · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.634747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.179341Z digest=sha256:35200d49e6de7d312088f255cee20b02efc56437584bb9b7de768afc5759cef8

Observation a39cf59c-84a4-4e32-b488-bc1979a61e5b · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.332892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.359642Z digest=sha256:9c47b2e819215ba69d6255c7ab1f8fc4759a819698f9ac37aa9011f8bb32d316

Observation 6357baba-1f9d-446b-aecb-937f7b65bf4e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:58.542592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:58.542592Z digest=sha256:ad2e73ae13a8cbf300b083d58ce5609a9ea3963336336b6674be08200f2263b7

Observation c60b0e3d-ffb6-432d-806c-cc2f525b1e2e · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.046429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.696284Z digest=sha256:26e97610bb03a12ed09ffccaa75de7585ed7ec87a77ce8d8729e5f83d494ba8d

Observation b861883c-b8da-4abf-abeb-f1d59ec23138 · outbound

This paper cites N.; Kaiser, L.; and Polosukhin, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech N.; Kaiser, L.; and Polosukhin, I

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:02.851166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.874889Z digest=sha256:98d65c633f1a29f28b215c74dd423db598d19c8623bfe76e3920cf37ce1a211c

Observation 7cf52851-2267-4464-ac50-5ebe5ba8e95d · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.042807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.042807Z digest=sha256:b864c06b09ba79a52bacc0e72f4813cb6d5e73df80fe228010b6a3a444ab6936

Observation 04e4ef9c-3e63-4826-b1d8-4ba7a760e217 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.231724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.231724Z digest=sha256:36de0cbc9c700a5c108b5cfc134100095f8de0f9e77acca646fd38dce0c6bd5a

Observation 87488ee5-d784-4d51-ba47-6cc934239f83 · outbound

This paper cites Qwen3 Technical Report.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Qwen3 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.316444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.316444Z digest=sha256:9dabf7a4ae4d60152ebd8cecb88c69254791a17386d8b0198c8bd72cd8341de4

Observation 3532d1ea-3948-4fe4-a0b4-111f73970903 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.444748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.444748Z digest=sha256:45cbc24d37e3443df19ce9a05c4d5b5bb55b8fca2ae9edf17929e9d2c9779f6f

Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · outbound

This paper cites Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.575812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.575812Z digest=sha256:946f118541a22a22726a8d092165d6264c2cd24d19e35b7263b93e8725c3ae6f

Observation 4b14c852-8f34-4761-bb46-b3461b79d073 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:02.663861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.683308Z digest=sha256:2336d6f1a1251b6dfc74aa6eaf049555769a660e6aae3b3082e7e905b5b42818

Observation 63838768-dc09-4026-8baa-6769dd2aae4b · outbound

This paper cites W.; and Li, H.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; and Li, H

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:02.432658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.768998Z digest=sha256:5c0eb2bfe2e99eb220bcef135a71380fef7493b352d729733121b0ec958896b6

Observation d5da1b1c-023a-4ea3-99d9-63d9d35a2d6f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:02.151155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.914789Z digest=sha256:f2cc676c63bc22334baf0decc81c71b5b1234efcc34cbf8ed4a248d36554b680

Observation 6cb74769-28fd-4cfb-baad-e0f37447b91e · outbound

This paper cites , " * write output.state after.block = add.period write newline.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech , " * write output.state after.block = add.period write newline

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:22:00.026278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:22:00.026278Z digest=sha256:4ebd57cb141a405b5d240761849d7301a41331a81419f2fb45a7834765f7f24b

Observation 4c4a7b8d-536e-41d3-aa1d-07323ea46f4d · outbound

This paper cites write newline.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:22:00.140173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:22:00.140173Z digest=sha256:747680bab57abaac123bddd212f7c08639c4bea2c0601c2a067b1688472d5c4f

Pith citing papers

Observation cf142752-2af7-45bc-9cb5-7725da8b0756 · inbound

TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis cites this paper.

TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-21T15:34:15.080474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T15:33:08.885667Z digest=sha256:14799503979f781606709c8e348a9cd2ea0390ee4f78c57909fea440646972ff

Observation e574c656-9807-4a67-97f2-0ac729782f23 · inbound

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion cites this paper.

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:57:42.883160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:55:32.044506Z digest=sha256:f112296755887227136f25869160a0d3567807295d2f40544d607c264837f283

Observation 3bf6e419-1013-4bed-8f91-02ff709e7272 · inbound

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios cites this paper.

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:21.928324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:21.928324Z digest=sha256:3ed3b38150d2ce3219b16cc3d3af33133751e7509206063bd977c28b9f4c7230

Observation 878d1998-5126-475c-b0c5-ebb94ffe3517 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.931802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:2512750a33d62148a741d7ec7870cb43607d52c62f30df4a89e80eaac2b7f636

Observation 5d1be585-afb8-4162-8a6d-24ba3001dc1f · inbound

Sharp spectral estimates for free boundary problems arising in plasma physics cites this paper.

Sharp spectral estimates for free boundary problems arising in plasma physics IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T19:56:33.793167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:56:33.793167Z digest=sha256:565b8355d6b2b8e3f72e56d5c23cd978bfdd399ac182ec13101018b5b9831a3d

Observation 3b01ef42-a07b-480a-a239-4229b954d18f · inbound

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection cites this paper.

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.994347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:55:09.141906Z digest=sha256:d0aa9388db6ce899fae913b203cc87e4c9b691205b124c5b96d91c31330e77df

Observation 0ef942c6-ff61-45d8-b44d-fa72ec1c2ae4 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.305794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:b0450f9986532a0bf4b48f9a6359da184310d17732aea33a2074ec147afa1181

Observation 713eefad-13f9-4a93-b473-955a704b9278 · inbound

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing cites this paper.

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:51:00.754781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:46:58.010112Z digest=sha256:541bc3b2ade91a801eec875529e57b940e4a884ffecc5f00cb7d2e62ea40cbdf

Observation 03223873-a2f7-4fdd-b941-df791230ca24 · inbound

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation cites this paper.

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.500305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:36:07.090627Z digest=sha256:388f215edeb9ca94ce81f72398308070cfe1922e445a737c3685dde5ff7a4f0a

Observation 8a395226-b98f-488a-bfcc-795612a049bd · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:10.007737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:601069d0d78dc7e692dc98debf36255cea5e2c15aecaba4d7d2413450bf07d0a

Observation ec3507fd-42e6-4b9f-af66-555e92b2e64e · inbound

RTCFake: Speech Deepfake Detection in Real-Time Communication cites this paper.

RTCFake: Speech Deepfake Detection in Real-Time Communication IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:26:17.854987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T05:29:45.895667Z digest=sha256:83ef7b0330e04c3b79acf4f98566ee656e4b53a292baf850868bec1151566ef7

Observation 598b00b6-b804-4a36-8a51-1ca6d28b3aee · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.102735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:3375189a465602a1e9cd13bfe553e84659fe6d75483e460e6bc991b5759ce444

Observation 3ba17a93-84ad-456e-98e6-97359ecd5833 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.806423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.806423Z digest=sha256:505ca41f04d423ac1318d59cad6fcfee93eb719bbb4b3d54ed720487724d7c84

Observation 235e690a-a3f6-4cf0-b49c-a4865a81eac1 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.070996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:3e29b7920410185f252be4b04f8cadb6e3884503f2ead2e4a5c80f6051893ae1

Observation fe258340-112e-4730-990b-7bc43fc0e81f · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.280453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:06:54.972846Z digest=sha256:0c2c465d902c6a1b6ed849296dbe0be4e7a53cc186b9f2f2dac1b98fabeb5f6b

Observation c7b3091e-e029-4632-9041-6273c134b334 · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:29:52.767862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:28:21.790110Z digest=sha256:f0e73cf4c2e1a0aba17bab9343ce574d7688141e9c3d0bf95054f9dfcaa79f1a

Observation c36b7f05-03d9-4e50-a73a-d1e5dccec9b1 · inbound

DeepSlide: From Artifacts to Presentation Delivery cites this paper.

DeepSlide: From Artifacts to Presentation Delivery IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:03:10.135599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T18:03:02.253422Z digest=sha256:d107de4ceb6e939fab2da9b4985bbc76269d931e907726fa6135e3eacaee0668

Observation cdc9418e-d138-41ec-8f3e-14063841f124 · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.031428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:87ef2794ca16d0419c86b1e0b3e2c6f1e7c4a74402fb3854beb640ef5a4054c9

Observation a5dc0fd2-ca5e-4249-aa34-398776f16c9a · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.232955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:78186b3d257abe2b1e04c9e3ce3466a74f3f9514798344bfd594e138f9cda945

Observation 579f3555-339b-4a1c-9417-fdf45cda185a · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:00:59.040259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:4f0ad26d9e1778fc96603ea9b96521ffcfbdb53ad75102c05ecff9131f97e36f

Observation 8144dd9c-78d0-4466-b216-679c9b414ff1 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.802184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.802184Z digest=sha256:5547c3015d7d8bf39a813b1807cef3d80adba4b1772af9c32eb79790711acc94

Observation 2830a0fb-9f2d-4eea-aede-ebac1bee36ff · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.082455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:f22c627cb2401071edbb98c2f2e74b913121aa7767a9a44d6bd2ef319fcd4d6e

Observation 22e6c8c1-4c35-4e15-8038-3119b51f83a9 · inbound

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech cites this paper.

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.523055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:35:06.352720Z digest=sha256:12e044b00085fa25c3f4d536010b922c9b1665e0736fd1da1957f92c207a416a

Observation 6815c925-6871-483f-bd92-24d4edcbc89c · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 107

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T20:46:13.835687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:9ced8ef677b9c705622e6317c9e8e780016a1a61a0f0aab7d2382785a7f0f35a

Observation 16f5240f-3f96-4f1c-bfe6-8bd5849f9f77 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:46:24.634388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:ffb05e92bc0f5a9a8b595bb28d34d942dee5e3559ff77d5595d01eb8180294fa

Observation 61fcceb1-562f-404d-837d-ba92ee5d31c3 · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.224385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:611098d9aff20b05cbba15fd07a159ddbe1fe8782db2bdbdbc2b4bedbf614866

Observation 0be20842-bbf1-47f6-9ecb-4f2e9eb7446c · inbound

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models cites this paper.

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:20.002939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:16:07.007759Z digest=sha256:a241e5f2347a4230ce11a5af33625ad7a4add01ab17da0015435d84c7d9c1d25

Observation cc9f9742-5fa6-4233-a3fb-1f7dd5b91dd0 · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:07:35.857767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:bf94139f2520d9be0feba64e881d4939252b4c3c2015d7397732c6e4cd626396

Observation 8581762d-eee4-4332-9ce6-51353a517a4c · inbound

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations cites this paper.

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.371843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:38:51.371500Z digest=sha256:d348c1d57643d9532348db67200a140c0901bc611efbc642c4d377424a9cb44d

Observation c83c0c63-cb90-42d8-827d-dff5dd9adeb3 · inbound

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction cites this paper.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.958514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:a1fc50715f3d8a82297d08fa7016e99504c1cbb187431b7459e0f5dc53ecfb59

Observation 7a8912fd-aeef-48a7-b060-41486a59264a · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.572183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:d483afed54563fcb40e628337f7fae58015d42d157b55db1b030505912b73b3a

Observation 3fcbdf0b-04d4-4188-a95c-e24b4faa8ff1 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.701979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:5dbc3cd27f3b1258e154f0614f90b08fe3e5be57ccdb59911dba2227ca6112f5

Observation 8dde18c9-d7bb-4c47-b60f-d6c59ab96167 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.634822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:837ab82501e2d1305b92d1cfd99c06836054b9c04f1ee33b3a259ca6427710e8

Observation 04d45292-d23e-45b6-b3ed-a2e00b37f448 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:1c1f0bbc69e4a717019e7572049969db90b88ffe32dbe7bd78d5a4ca659a3108

Observation 997da372-83b8-4418-a395-df4cb3399f5d · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:54:40.709861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:35:03.660671Z digest=sha256:99c3d2927180b9f822f14596f34a4b98cd69d2ed1362ad40deb8bd06d1762943

Observation 5b46fa76-577d-4a5a-aaf2-9b813acae9fd · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T11:11:08.878850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:11:08.878850Z digest=sha256:081f40607ae438d615ed8798ab2d74e74c5363d746c905b7941444773cd21c64

Observation cf0a7c94-810c-486b-80c0-50d42fcc5bd4 · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.182730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:415a126f95bd706d13655704eb5b26e007e402c342a6d37463554ff164266665

Observation 28686fc6-7cc7-4693-887f-911c13a40eba · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.726590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.726590Z digest=sha256:2290af97a8ad3ebe1e2fd024210204647fb6966fab7e22d52bfc96aa910a4bdd

Observation 8a132f17-e405-42be-b8b7-27ec31c2d703 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:31.531580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:31.531580Z digest=sha256:d0dbb49c92c272f3ba1018184d3f3f6af8648916f895450ef30ec2b831498824

Observation a0cdb54b-a9a6-4ab2-9ced-30c4f7c6082c · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:50.544402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:50.544402Z digest=sha256:d021748337702e8ff7d0eb541806a9cec130aa23eed3249c549e2aec7cf6382e

Observation f2a5988c-5344-4936-b561-59216c21ed4d · inbound

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models cites this paper.

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:39.932052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:39.932052Z digest=sha256:419b135428f34c35e124d1c8f7a940a4736429fdb9df3175903449d3d8b062e0