Pith. sign in

Paper Citation Record · LEDGER

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 44 inbound Pith citation observations for arXiv:2506.21619.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21619 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:22:00.140173Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:40:24.981589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:19:49.633371Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba6482a5-3d62-479f-b595-93150211c630 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:51.474748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:51.474748Z digest=sha256:890c94740cc5feee1d7fae5451b554d4178b974ff150d53f06088f696a13ce81

Observation 1fbb62bb-fe07-419b-8184-9c3d8349b7d2 · outbound

This paper cites C.; Vidler, J.; and Roedig, U.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech C.; Vidler, J.; and Roedig, U

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:09.568747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.714747Z digest=sha256:06e4371e17dd03d53f974a6063cc61ccdfddf903f2b1c13668f260da4603c8f1

Observation 857060fd-8691-4b3e-a815-b652d61d99d7 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:09.343035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.891560Z digest=sha256:a5eafeec8a844b44b86868d139c809c86a5935daa23eecfc2e4c754ba6c29e64

Observation 55ae0086-9ecf-43fc-b66e-0443261bdd61 · outbound

This paper cites o lge, E.; G \.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech o lge, E.; G \

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:09.141449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:51.984742Z digest=sha256:3c6c98556a8ebe634a0565855e5a70d7a964a5dfa78b3e577a87b4283c02b484

Observation ed11d40d-2d66-4e37-8525-c8ea8db6105e · outbound

This paper cites T.; Rubanova, Y.; Bettencourt, J.; and Duvenaud, D.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech T.; Rubanova, Y.; Bettencourt, J.; and Duvenaud, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:08.907713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.090822Z digest=sha256:8636f7b9618537ec5f3e09fbc6a5361cc53dd48df09c15a71e6ae37d0a933b3f

Observation 7c44ba08-afc8-4dbf-94ea-b407c5446dd0 · outbound

This paper cites Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.191734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.191734Z digest=sha256:a03158732e684c86e29417c4d049ef20a5c3ae5bee3b424fa1fbea73064a1aff

Observation 1e47c365-9195-4e19-8118-c14e329e8df8 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.664334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.374989Z digest=sha256:99563956ae2518581f5a9805edf207a9eb60f1c78eb02a86dbaeb0ce5fb05d30

Observation c7f49650-edd8-42f0-9b75-ed6c67c51a73 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.491309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.491309Z digest=sha256:92a1967691561bd999c936bdcf7d1702d08904ac7d8664d7b5db0c9cdf3540ad

Observation 3389ed9b-60d1-4229-a949-802ace936876 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.443414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.584839Z digest=sha256:e4b1c1474c3fa2e0eeb21e10e417fe9dba54e442bbf6190555a6b66c524937ae

Observation ed3fd843-b8d4-447d-8fe2-389ea4370182 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:08.224525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:52.694748Z digest=sha256:e62f9a7219c002fe05405f2c925e6e489a2839e52e958c0d32a179d10b4e5bfa

Observation 743b8f20-f8f9-484d-9729-16dcc0f97530 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.901086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.901086Z digest=sha256:d157534d23208c630a5e4e77f62f0a8a2aa27f0ed151a6babc87029f939dc1f7

Observation 2e9b7afa-6e45-4de3-965f-74d7e921dc57 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:07.977427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.044748Z digest=sha256:dcfbc6611bae3baa078cdca8d01529228000ec3b2c558f81913806ebe95b00c1

Observation eef22172-612c-406c-8739-4ca04512428d · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.174876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.174876Z digest=sha256:87149703b5d2fe6cd31503ac3d7b093fcef98bd81da595d761bc3370de8398eb

Observation 109b26b9-bb32-4ca1-a88b-3a458c154038 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.286323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.286323Z digest=sha256:1d3c18a4780a7a9ced359cc5c7df46abb0458cae118118bf02d0ea227dc58efc

Observation 248496b5-51e7-4c45-834d-fdca8bf8d666 · outbound

This paper cites A.; and Wang, H.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; and Wang, H

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:07.662270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.404945Z digest=sha256:e8ff702c40c1972915570e5dbb51917a8542b548c17aced8fa5b6832bf2e5f26

Observation 68dc8c7f-92f3-436d-8588-5b356d0e5ff6 · outbound

This paper cites E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; Liu, Y.; Zhao, S.; and Kanda, N.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; Liu, Y.; Zhao, S.; and Kanda, N

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:07.397033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.514918Z digest=sha256:9b41d4f4d574344329f30ab8a636a0efaa1c60dd1f48db090837b29876866a8a

Observation ff85a550-8a16-46de-9218-987d9ba51989 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:07.073937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.654824Z digest=sha256:78991c428b9c7f09b45d2396253d31f6ec7d356d62edeee955a26afa8aa01d93

Observation 3ecb148e-6917-4155-b0bc-9b5bbfe8e867 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.782697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:53.794860Z digest=sha256:41b117a22b4eb307a5da6446c1559b90f6b537dafdeebff2ebc2ad7c4431778d

Observation 1112caf4-2c01-4233-bab7-13bf90d6c3d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:53.970330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:53.970330Z digest=sha256:71da52aad701097b0a5fdbb2d13f7a44ed2f8e6268e73ff3d6c0afc2926e4824

Observation f04ef65d-a581-47f3-8108-7d8574e282b2 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:54.069634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:54.069634Z digest=sha256:feede39a767c4901f1e6e9ed9480f07aa0ed3d28823e275c23a27ca8a3d356c5

Observation c8b94280-cdb3-4f31-a909-433f45771f95 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.449439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:54.324773Z digest=sha256:3d5f9e527d87413b96e1704cccb3e56f7ee4fa0fb2d1976f512102ee2fcdc30e

Observation e314fc3a-0310-4ba0-888d-305ab9815519 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:54.509932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:54.509932Z digest=sha256:6a0b93b1130ece2229e45da88d4e5ee01f4752c66ec4c0d63891ead889e7d586

Observation 15ff03d1-ba1d-4f90-b261-66173c4e3946 · outbound

This paper cites ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:22:01.640762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:54.594770Z digest=sha256:43d11f61df4abae2e4dd18a51cad6a0175bdeb7a136bf6f908c5b32654cd0bd2

Observation a83a6da2-1f0f-4950-84ee-9026ec8c28eb · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.093152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.093152Z digest=sha256:b78e73fef4e7683026dbd8946653c0ac148d171ad4b217f70b2ca81ebc3cf544

Observation a2e58924-eafc-479b-b71f-e7919cd052b9 · outbound

This paper cites SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.224747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.224747Z digest=sha256:aacee0d5693b50f46c16541df7e48ab149aadca3ef370cee91b3ce3cc93871e7

Observation 50297084-4b08-47df-8773-dffb35428c2b · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:06.135488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.490696Z digest=sha256:c411c9de43488fefcecd3970e3f85c0ca69edc7039a35c98ef55e9f9794ca11d

Observation eb32930b-3b78-44ad-baa0-fadb9101c5b4 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.968350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.624763Z digest=sha256:58402f6adec7b07c031edce8f91a9c8e17da5e74958bc3265f9d9ac17a5409d3

Observation 0fe15248-fe58-4e00-8cc3-3c21292f24be · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.719910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.719910Z digest=sha256:78dc40058becfd90f3a91689974045e9d5105b9594777ed6164eb61e93ba6442

Observation 76fed29e-242d-4dfa-b6fc-41b2f8fc9140 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.843335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.843335Z digest=sha256:179aa1e197ac8ba7cbfcb978ee47a52271ccdb55c6eab767a66aaacaea9df702

Observation 55540a81-a7ae-44cd-91a1-fd5e4d506dc2 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.796837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:55.959394Z digest=sha256:a902178cb3baa44c1dd3cef60821d98057bda98163ca92f6b5b806faf6deafd8

Observation ec264d95-654f-440a-9c15-2b2f25920f8a · outbound

This paper cites FleSpeech: Flexibly Controllable Speech Generation with Various Prompts.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.065145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.065145Z digest=sha256:c6588ad03ca18a37907a6b696cf62d06b86b0744dce200452109287f0a839e36

Observation 2b651a00-37be-422f-bbfc-7bbf58c16c60 · outbound

This paper cites A.; Han, C.; Raghavan, V.; Mischler, G.; and Mesgarani, N.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; Han, C.; Raghavan, V.; Mischler, G.; and Mesgarani, N

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:05.560496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.158020Z digest=sha256:e5177cf28294565815de129ebec2714dd820012a12989f647209aee6af9a8916

Observation 2bdd39a2-b877-43b9-ba76-26e40715669f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:05.244739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.230762Z digest=sha256:c3081175bb5fb3e554d45a238d44d13b60d58db97577256748eaf603991649e7

Observation d4926be1-b5a6-4c69-9098-2f81b5e375b4 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Zero-shot Voice Conversion with Diffusion Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.314992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.314992Z digest=sha256:92553a81c5d62fa896d4f8ac97d7ada45af72033323e724b95ed5a8fd89fb753

Observation db8cd0e8-bcf4-4dc7-9249-62ea0776d5a3 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.820065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:56.474829Z digest=sha256:58a788693c142679e92e627771a5169bb427c33d71f56128c3dc986c7e45828c

Observation d7051c0b-f641-4363-940d-13858c01e34d · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Finite Scalar Quantization: VQ-VAE Made Simple

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.557191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.557191Z digest=sha256:441078104c27fd879bfc217ddd320423dfb6e2bb8835af978f92e13a34fe552b

Observation 25004936-e1d2-498e-a7ea-46875215afb1 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.630048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.630048Z digest=sha256:bf8ce2e8909c1a934c4b561deac7763e3f1dd79c43f48eb15b58f8acfad478f4

Observation 42046312-291d-4ca5-8cb5-79b7fa1cdd94 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.720288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.720288Z digest=sha256:0a131672352937978fa38d09bca2cf22263b6b495192c72201fd0f3a01a4e994

Observation ca11a80b-9add-4a97-88fa-d01d3087f14f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.538604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.085294Z digest=sha256:0b165aebe76160c119f0751e74fdcc87acec00c9edcebb845f6b52a33ff12d1e

Observation 51dd94e3-a1a6-48c6-9188-206b659b9008 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:04.386285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.212866Z digest=sha256:4ceda09386d465e83db24a6067e5807744658dd2f163d4f589441dcdc70e9a95

Observation 9ecc25f5-0aeb-4093-ba3e-7273c48a84ff · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:57.367656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:57.367656Z digest=sha256:a174c4e054f0d24d4df094597df1d25f9e4c05563598f088f7a228158b28b41a

Observation 4753148e-6bf7-473f-8d88-4f3c673be2e8 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:04.191326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.484383Z digest=sha256:68a156da2648865ac7f9f558c7b1f0aeb4e29f3773bf957988015d2fb2521e28

Observation b6bf2331-1b89-4039-938e-dc3b51fde4bb · outbound

This paper cites A.; Gonzalez, J.; and Escalera, S.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech A.; Gonzalez, J.; and Escalera, S

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:04.028061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.624749Z digest=sha256:0966d393760ec821ae4583f49bb8d63c9b87a6c21f247415d86f43a06a3c56ac

Observation f6c00cfd-550f-4e87-9953-321632221f2c · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:57.740164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:57.740164Z digest=sha256:d3799ea54914782e0471eeadf2e5f98dc9078c0b4bc3af4d367b089b2230bf30

Observation d58f3e58-36fc-4fd0-8dcd-579b40551c1b · outbound

This paper cites E.; Hinton, G.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech E.; Hinton, G

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:03.851801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:57.864744Z digest=sha256:9e58501aeec68dc63c15155525c1c416944bb2c2b2d11eef6b22c4ee0f624c75

Observation 214db0ad-301f-4718-9fc1-d18dda973ad7 · outbound

This paper cites DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:22:01.078010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.044747Z digest=sha256:780d6a22cf8e7603f7221c0a6476cce816076ea9b13c0e6fdc67307661dc4f15

Observation e3098054-309b-4bc4-bf24-f287f3b7d03d · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.634747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.179341Z digest=sha256:c28f5bd35d13ceb88027729fd52c9498572ec115ca88713262c7abebb2ff798d

Observation a39cf59c-84a4-4e32-b488-bc1979a61e5b · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.332892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.359642Z digest=sha256:688f7e4073b64e97749859cf62398c7d9a9f25a8567fe96d6f07377dd6291f27

Observation 6357baba-1f9d-446b-aecb-937f7b65bf4e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:58.542592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:58.542592Z digest=sha256:f84b697cb5c723e4da970464c2b114506be9fb39d8ead6edf239073502865bb7

Observation c60b0e3d-ffb6-432d-806c-cc2f525b1e2e · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:03.046429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.696284Z digest=sha256:c55db8b7dc1be06faa349335811711225ed631d9d815f92663aa11d0c4f20606

Observation b861883c-b8da-4abf-abeb-f1d59ec23138 · outbound

This paper cites N.; Kaiser, L.; and Polosukhin, I.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech N.; Kaiser, L.; and Polosukhin, I

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:02.851166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:58.874889Z digest=sha256:6ecd697e2994560755b68784f6beae5abdabfd0922fa3dfb53c9772b7e2d0163

Observation 7cf52851-2267-4464-ac50-5ebe5ba8e95d · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.042807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.042807Z digest=sha256:640ad50dc2f10c14908d7437f04fc995c89460c1c2f4b706260efb99f421687b

Observation 04e4ef9c-3e63-4826-b1d8-4ba7a760e217 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.231724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.231724Z digest=sha256:f547fb87361e1a1a23e8ffe2bb2d777443d498e5059f8f9c5d95ff63fadc8c28

Observation 87488ee5-d784-4d51-ba47-6cc934239f83 · outbound

This paper cites Qwen3 Technical Report.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Qwen3 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.316444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.316444Z digest=sha256:144f764fbb0ea6b68e058c8a7f611aea2bc4f9ab6de6b60fe5e0519ca93bc69e

Observation 3532d1ea-3948-4fe4-a0b4-111f73970903 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.444748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.444748Z digest=sha256:26b87fa3309fc1c3044bb84968453cac973e766b2704e9577277254690576956

Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · outbound

This paper cites Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.575812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.575812Z digest=sha256:06494f8b5da9fff2688be221461baf9f74ab55b371036db9d89b245f384f17c4

Observation 4b14c852-8f34-4761-bb46-b3461b79d073 · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:02.663861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.683308Z digest=sha256:13aa01f76836df0b05afa807067b0fa214c1fb239497310101515877d8eaba01

Observation 63838768-dc09-4026-8baa-6769dd2aae4b · outbound

This paper cites W.; and Li, H.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech W.; and Li, H

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:22:02.432658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.768998Z digest=sha256:c66d82da1dc36add787ced1f1293a04de8aba71082545090f62bec6ba35b5131

Observation d5da1b1c-023a-4ea3-99d9-63d9d35a2d6f · outbound

This paper cites an unresolved cited work.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:22:02.151155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:21:59.914789Z digest=sha256:796838fa37971af6400fda949b7a7a5828f19205371293be3d51a99c1bfdbdd6

Observation 6cb74769-28fd-4cfb-baad-e0f37447b91e · outbound

This paper cites , " * write output.state after.block = add.period write newline.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech , " * write output.state after.block = add.period write newline

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:22:00.026278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:22:00.026278Z digest=sha256:e1c31e5626c3cac1f13efe9881d34484f3fa46e6214c9dfe942f81a43a300a69

Observation 4c4a7b8d-536e-41d3-aa1d-07323ea46f4d · outbound

This paper cites write newline.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:22:00.140173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:22:00.140173Z digest=sha256:bff63b5452ceb603fd973efb5cc4e349a7e829af3ea2ae81f5e8a98950463474

Pith citing papers

Observation cf142752-2af7-45bc-9cb5-7725da8b0756 · inbound

TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis cites this paper.

TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-21T15:34:15.080474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T15:33:08.885667Z digest=sha256:4de1426a1bdce6f796866bff19ed71d6e3de68f3f2445cd3b756acb17b35424b

Observation e574c656-9807-4a67-97f2-0ac729782f23 · inbound

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion cites this paper.

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:57:42.883160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T09:55:32.044506Z digest=sha256:e2ddc58872cab74eddcd72d4603dd0669013c259b83919fdb436936402b95547

Observation 3bf6e419-1013-4bed-8f91-02ff709e7272 · inbound

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios cites this paper.

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:21.928324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:21.928324Z digest=sha256:99a07a7b886a429f8553e15cce1e3d71840b5fb7fd5e134741491eb1b0d772e8

Observation 878d1998-5126-475c-b0c5-ebb94ffe3517 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.931802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:8ccd1e8460817b7bc92ccd267bddfc43210a384b7baed10f1b5196c2ad49cf4a

Observation 5d1be585-afb8-4162-8a6d-24ba3001dc1f · inbound

Sharp spectral estimates for free boundary problems arising in plasma physics cites this paper.

Sharp spectral estimates for free boundary problems arising in plasma physics IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T19:56:33.793167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:56:33.793167Z digest=sha256:7f51461677c258a2291662e7563eb0e449b5ef5a62f23d27e8ef79a5a9e9dca0

Observation 3b01ef42-a07b-480a-a239-4229b954d18f · inbound

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection cites this paper.

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.994347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:55:09.141906Z digest=sha256:55561d2afffbba324ec163e9ec85be1d9a36c498e4aa95cdc318ba4021c51293

Observation 0ef942c6-ff61-45d8-b44d-fa72ec1c2ae4 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.305794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:25c4bbf60e19fa1ea4c9718225631f2b60a10104dbbf2ea5d76ca5beab27945a

Observation 713eefad-13f9-4a93-b473-955a704b9278 · inbound

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing cites this paper.

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:51:00.754781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:46:58.010112Z digest=sha256:3b026f8590fa89cabb8d0d2f92d34d3cf53cccff80d1741c65971c61f8636ff6

Observation 03223873-a2f7-4fdd-b941-df791230ca24 · inbound

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation cites this paper.

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.500305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:36:07.090627Z digest=sha256:b75751d8156159ea212a139c54f7ef0264a1fb9668a4f1a467753799c8c02727

Observation 8a395226-b98f-488a-bfcc-795612a049bd · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:10.007737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:9486fe8bbb464a087bf1f22ea7ad09115b321916eedd97bd54b178925d239f58

Observation ec3507fd-42e6-4b9f-af66-555e92b2e64e · inbound

RTCFake: Speech Deepfake Detection in Real-Time Communication cites this paper.

RTCFake: Speech Deepfake Detection in Real-Time Communication IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:26:17.854987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T05:29:45.895667Z digest=sha256:870afd0cad5a0d04ca90c4c49247109297ed7c7edceb2fbb6c406737c91858ef

Observation 598b00b6-b804-4a36-8a51-1ca6d28b3aee · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.102735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:a452fd6d71ae4d3212edc7279341d40f009d5211f3ef87120d496fc88f611114

Observation 3ba17a93-84ad-456e-98e6-97359ecd5833 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.806423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.806423Z digest=sha256:59af80f852e94893d6d3fca72f5b1fa24b639812355390b84a6bef8877b5474b

Observation 235e690a-a3f6-4cf0-b49c-a4865a81eac1 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.070996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:beefe4887236dbee579bbceb5dd3a3d231945922724ad0b84f2513eb98815b8c

Observation fe258340-112e-4730-990b-7bc43fc0e81f · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.280453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T05:06:54.972846Z digest=sha256:285aa31ff8654baca80f7f0d8272443ebb54fa9e7536cc256f1fc34103980267

Observation c7b3091e-e029-4632-9041-6273c134b334 · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:29:52.767862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T08:28:21.790110Z digest=sha256:42e095b00928f7d7c10ccf398351557a13bc54091cddbbce9d2bcbfdb8c1c609

Observation c36b7f05-03d9-4e50-a73a-d1e5dccec9b1 · inbound

DeepSlide: From Artifacts to Presentation Delivery cites this paper.

DeepSlide: From Artifacts to Presentation Delivery IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:03:10.135599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T18:03:02.253422Z digest=sha256:eb120aa83088f55fb1c9807120ed0b17044447aa5762544d1bd246c26e80b014

Observation cdc9418e-d138-41ec-8f3e-14063841f124 · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.031428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:df6c0db99d5e59bc10cba7229d294dadd3c6c57fc3c15a61b82ddeece98ee8cc

Observation a5dc0fd2-ca5e-4249-aa34-398776f16c9a · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.232955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:ed7cd09bf7d03d012c4147604cea11a5d268a3090263181bd8cb29171467c640

Observation 579f3555-339b-4a1c-9417-fdf45cda185a · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:00:59.040259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:dbf1340ff66a44703c32bff7b121fd960e13f90c3cafec3c79dd81aac5d1826a

Observation 8144dd9c-78d0-4466-b216-679c9b414ff1 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.802184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.802184Z digest=sha256:738a8bc06ac23fea018dfb0b8bc98e38f98568bcddea4357bc28186f0532df63

Observation 2830a0fb-9f2d-4eea-aede-ebac1bee36ff · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.082455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:c4f8a183a94bbb5e9a7d4a940aa60b1c23719f171310236003fccef75836cb27

Observation 22e6c8c1-4c35-4e15-8038-3119b51f83a9 · inbound

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech cites this paper.

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.523055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T17:35:06.352720Z digest=sha256:b7821f149e0f2d6447d9dc97f5175d275c6d78f3fd286d56e5a37fe4f98fb9db

Observation 6815c925-6871-483f-bd92-24d4edcbc89c · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 107

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T20:46:13.835687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:d2b393b3425bd8f0c4191c4259c2fdee6ab7cef4bb52b23281cb02e02294dcfa

Observation 16f5240f-3f96-4f1c-bfe6-8bd5849f9f77 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:46:24.634388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:09332c0cb1b0d69adda10a8f19dc67251f0bae8a9bbc37a392f8ff0f32c9dd5e

Observation 61fcceb1-562f-404d-837d-ba92ee5d31c3 · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.224385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:cb3fe7c069ec3881fa0c6f31c86fb0e71bb06cbc2c4b017ab801564efe80b123

Observation 0be20842-bbf1-47f6-9ecb-4f2e9eb7446c · inbound

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models cites this paper.

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:20.002939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:16:07.007759Z digest=sha256:baf71a2304e496932cbe05accd04836e2f44826871f273bc5681c19e500b20b0

Observation cc9f9742-5fa6-4233-a3fb-1f7dd5b91dd0 · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:07:35.857767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:362d5b7d3038c70ad4c26bfbf4e627a46f8fd03a17a55994dac1e857dd2e8928

Observation 8581762d-eee4-4332-9ce6-51353a517a4c · inbound

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations cites this paper.

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.371843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:38:51.371500Z digest=sha256:bde34c556eefa20129d80f6aa7d2ac4288f1a48bb3720b080851748be6297bd8

Observation c83c0c63-cb90-42d8-827d-dff5dd9adeb3 · inbound

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction cites this paper.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.958514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:c9bc24827c468892cc9080cd5ffaecc2dd4b97d8bc17a95674a3a0c5df58466e

Observation 7a8912fd-aeef-48a7-b060-41486a59264a · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.572183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:91999f123420c8a21a1d2ddf799cf711f956439c155b8b4bde67df5faddaf0ff

Observation 3fcbdf0b-04d4-4188-a95c-e24b4faa8ff1 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.701979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:4c2db405f73567b0ca06142ee241a03febdcd471ee9fcd5519b09ddeb9e5a8f3

Observation 8dde18c9-d7bb-4c47-b60f-d6c59ab96167 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.634822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:2cc944353d1d8fdab6b3681ac04b91abefb18c74f1be957b9aa11bfbcf2fa19c

Observation 04d45292-d23e-45b6-b3ed-a2e00b37f448 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:79d1f5a08935a9b2e5462c60546f048bb75994b23795e1f9eb4e62bdeb48b548

Observation 997da372-83b8-4418-a395-df4cb3399f5d · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:54:40.709861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T09:35:03.660671Z digest=sha256:898373c7200c64bc1b2784d73c78faf6d0450581965f1af2e72ff824d8be6d54

Observation 5b46fa76-577d-4a5a-aaf2-9b813acae9fd · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T11:11:08.878850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:11:08.878850Z digest=sha256:9b91267601ad67a90d3fcf60e75f9637af354ba8b9eb87b6d5a52e9eb155dac2

Observation cf0a7c94-810c-486b-80c0-50d42fcc5bd4 · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.182730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:9569d90a262d65041ca7e7d060328bd2af90df2439a8188736f3ba10ba426242

Observation 28686fc6-7cc7-4693-887f-911c13a40eba · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.726590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.726590Z digest=sha256:17d982b330631d319148b9f60092870727b230895d3826f80e07acce0f9c6062

Observation 8a132f17-e405-42be-b8b7-27ec31c2d703 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:31.531580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:31.531580Z digest=sha256:1bb41ef1bdba0e3556c0e26c519b15c38c3373386f98ab25f6280c966c48b9c5

Observation a0cdb54b-a9a6-4ab2-9ced-30c4f7c6082c · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:50.544402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:50.544402Z digest=sha256:d5763dca343b8f7a314d97f1aedf1a6e24cb787e95ba2fa206dfb54634c71dbb

Observation f2a5988c-5344-4936-b561-59216c21ed4d · inbound

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models cites this paper.

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:39.932052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:39.932052Z digest=sha256:56c92adf6df27e26bc30bf630752892ce193fa31f9ec98875b9f639559bddade

Observation becd0aca-8ccf-49e1-a195-9a55ee45444e · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:41.988791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:41.988791Z digest=sha256:4b7d0cd2bd63838b35011799ab9e952f3aa26c3799a0f6011673a13b03aadc9e

Observation 118e73e2-b728-42ec-ab50-4485a16f7c18 · inbound

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents cites this paper.

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:47.839144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:47.839144Z digest=sha256:64916528987b7501a37aa9657b9bb242b68c24a4e884f0af6ac8898175890b39

Observation 2a4a3f19-edd3-4042-8c13-52514f576082 · inbound

Luna-TTS Family Technical Report cites this paper.

Luna-TTS Family Technical Report IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.981589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.981589Z digest=sha256:b4a1cd6c5deddb094628fec47ef2342961b3d26d2f2c1a931798ee71f726541e