Pith. sign in

Paper Citation Record · LEDGER

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 19 inbound Pith citation observations for arXiv:2506.13053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13053 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:45:17.467537Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:40:25.263064Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:34.896349Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46dd6f12-388d-47f2-b4d5-1b58801c4773 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Neural codec language models are zero-shot text to speech synthesizers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:24.192871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:13.214738Z digest=sha256:7dbda23ea4e4b80926d11db0ce179c9df7aa463f7233480b2a86dad14c0e5ff9

Observation 2cbfe618-9751-4ea9-90ad-f4aac7513486 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.288815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.288815Z digest=sha256:ed81086da7843da894783e58ff7db82375dc1a52da67ebf87d419bce2ba4090a

Observation d4840c43-7c3b-4257-980b-e308b9d53508 · outbound

This paper cites E2 tts: Embarrassingly easy fully non- autoregressive zero-shot tts,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching E2 tts: Embarrassingly easy fully non- autoregressive zero-shot tts,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:24.038565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:13.380779Z digest=sha256:80bc260d8a2e670e5111b3d62c1973a14cb9e671fbe4853475c7c136876b43e1

Observation 19e19724-5546-4746-8948-d7be6e8427d4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.443312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.443312Z digest=sha256:6102485c43b53972b8e6c689e3af23e0dadbf83e54ba2f8850714c60717b49a5

Observation c4ee4589-42be-40e8-b106-1220f2f921ad · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.893891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:13.540878Z digest=sha256:4e8127747499cb04daac9743b8055f69513349d57df4f471d100942c45ddcacb

Observation 47614f9a-f3b0-4357-b723-32f725028c9f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.608161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.608161Z digest=sha256:97f7126e70405e526b9a694586222c9653db41894322c7c74de6b709d0d7d7fe

Observation 811b01f7-0773-4610-a07c-7bf31ea74375 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.675822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.675822Z digest=sha256:f85f6a8dd6044201e8d839630656c795becf2b56e5cc146ac6daf5dc22fe9de6

Observation 48f1086b-b0ef-4d23-aeb1-644adc8484d4 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Libritts: A corpus derived from librispeech for text-to-speech,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.776702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:13.755949Z digest=sha256:f4179fd59cb903410043b73f9f195a620b8b28bb460d5e89d516764464eaf429

Observation 1a7c5959-0e17-4458-b945-020459b91c96 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.847783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.847783Z digest=sha256:e5120c2f46e23de74d691dcc5b3c1ddb2fd6184a069949db9529780f6a131100

Observation e190dbb3-1fbf-4a1f-9e04-e730369071e3 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.630165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:13.919856Z digest=sha256:95cbfdd4d6dedbaf9b7332c59dc35e6854c524ae7fa61054a183371e2104ecbe

Observation c4b312cb-f7a2-492f-b2c2-3ea22dbe1a6a · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero- shot speech and singing synthesizers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Naturalspeech 2: Latent diffusion models are natural and zero- shot speech and singing synthesizers,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.479251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:13.988814Z digest=sha256:6e4ee5dcc784fa6402ef9a2480f16bc443dd9c1ae6c64f30dec5528945e17dc6

Observation 0648c2e9-ec80-40e7-b40c-854a16fed321 · outbound

This paper cites Flow matching for generative modeling,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow matching for generative modeling,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.333080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:14.064823Z digest=sha256:20fa702492e4a879438eaab08416873686c2ce01efbddf913cf4ee7756e1b6d3

Observation c726225b-2ddf-465f-9cf5-1cc4fdbf401a · outbound

This paper cites Sf- speech: Straightened flow for zero-shot voice clone,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Sf- speech: Straightened flow for zero-shot voice clone,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.200164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:14.143790Z digest=sha256:aa2112a26292247291d82d26c9bd15dfbf78793adf912b63e21aa36f47db7823

Observation f7b3f264-2b8f-494d-a920-a01aff7949ae · outbound

This paper cites P-flow: A fast and data-efficient zero-shot tts through speech prompting,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching P-flow: A fast and data-efficient zero-shot tts through speech prompting,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.080631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:14.242517Z digest=sha256:c98bec475f56c30803dd38270dc91b13b7076201144349d88c888adca5d2bde4

Observation 8e3f0334-850f-41da-9b83-dbe729364cf0 · outbound

This paper cites Attention is all you need,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Attention is all you need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.325866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.325866Z digest=sha256:8b037c0fb91c4093733106b25d84a42b6e5334e04af73012375397405b9b3394

Observation 1e0b8f32-b682-4c29-ba91-3741fc249547 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Zipformer: A faster and better encoder for automatic speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.897794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:14.417592Z digest=sha256:b206dbe3fb02082366ca87efb23a1fd744d2f8d3ad0badc974d45903ec6f0af7

Observation 2110306a-b124-47bc-be6c-8bd87d96343e · outbound

This paper cites Classifier-free diffusion guidance,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Classifier-free diffusion guidance,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.760185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:14.504825Z digest=sha256:30ef1b9ac7111c04070babc719ec7e115210ed497f6ed19277ca65cd306e3e2b

Observation 170f12b1-dcbb-4192-bdf7-bbe89d575b1c · outbound

This paper cites Freeu: Free lunch in diffusion u-net,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Freeu: Free lunch in diffusion u-net,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.605205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:14.572199Z digest=sha256:8f106c06630de7e3a5206c536ce64a2c7e80ab513b3cdadaa6ac5b35cb4a68bf

Observation 4a2bb12a-be56-4f7b-9cf2-5cc4415addbf · outbound

This paper cites U-dits: Downsample tokens in u-shaped diffusion transformers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching U-dits: Downsample tokens in u-shaped diffusion transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.210548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:14.658013Z digest=sha256:6b54cac6f1ab9e0b3abeba1c11f31a0752eee15199440fa08e831f504d121abc

Observation 1fad3369-6cbe-4ed7-be2d-5ae2ffe0c242 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Fastspeech: Fast, robust and controllable text to speech,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.731869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.731869Z digest=sha256:7fb71e08a20f06449e8ce8938c3e1acaacc28070736bec05dae2268c31805008

Observation 94ec5498-4dba-42f3-ae84-7ae06ad8cc7c · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Conformer: Convolution-augmented transformer for speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.047617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:14.832424Z digest=sha256:cbd2af9560ff0e64a8668861742d520eeda141686aafab2f0f45ba7ba4881715

Observation b9194bd8-2d98-4212-8bfe-fd2d1567bdf4 · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.930725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.930725Z digest=sha256:9650cfdd68a010c8928683e2a1c9f3e5b420be561175d30f89a305907eff410e

Observation 64bcbf28-fe5b-4eff-a317-75f4b2b627ca · outbound

This paper cites Flow-tts: A non-autoregressive network for text to speech based on flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow-tts: A non-autoregressive network for text to speech based on flow,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.877248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.033007Z digest=sha256:058d2bf59b3d6894e1bee75f0f96e5a6c43cc91ac442311f25dc9ea1a77cd4c9

Observation a84aa383-803f-4188-93d9-7e3c6359d2ab · outbound

This paper cites Simple- speech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Simple- speech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.707431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.151321Z digest=sha256:545ce1f9860719de5c70b0f9759b893da007035dac949a160f8bc5def93a97e0

Observation e5c11d3b-3dc1-4568-b894-f926fe03ec62 · outbound

This paper cites DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.441890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.249852Z digest=sha256:7c73b0db89981814ac30ee982344e2df2ac78990ecb8eeb51f228039693866fd

Observation 9954c653-8243-48f0-8f7c-bee76df3b892 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Convnext v2: Co-designing and scaling convnets with masked autoencoders,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:15.372934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:15.372934Z digest=sha256:385e9c77e7cc74f36f666d2c7cac75db04b7ff213d8b91fa0e2dbebc7370213f

Observation 38ae1325-a8b2-453f-a1ba-4bc4e81cc555 · outbound

This paper cites On distillation of guided diffusion models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching On distillation of guided diffusion models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.229679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.438075Z digest=sha256:1d08645780c5b81a427fc1723f339490dcf8edf16ebde494535d7b74581d30c6

Observation 8a9aeb93-b8ac-45a2-9e03-00166a481c25 · outbound

This paper cites Tacotron: Towards end-to- end speech synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Tacotron: Towards end-to- end speech synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.032849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.495205Z digest=sha256:15857ad940b6350606a98df7e07a45895802639d2c37f1b15f0703807de615bd

Observation c94e5b87-c82e-415e-b725-0232d8a39a93 · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:15.573777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:15.573777Z digest=sha256:9f38841b9beb9a6ca4b9625fee380af97ca913075c702a9caee916ffa31a316b

Observation 40a0df60-3838-450b-ae44-fdff04008690 · outbound

This paper cites Revisiting Over-Smoothness in Text to Speech.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Revisiting Over-Smoothness in Text to Speech

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:45:17.675731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.639463Z digest=sha256:8eb66a570ee9488abe96485e96904004b120cbf03f63fab39c021e88c9ace1aa

Observation e61555de-6d6c-4486-9be1-8c877b4ac97c · outbound

This paper cites Grad- tts: A diffusion probabilistic model for text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Grad- tts: A diffusion probabilistic model for text-to-speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.775881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.700325Z digest=sha256:f06493becddc6c8ef2331e38ea5ab0f59f55304f7c05daef6e8721e1a6dc3863

Observation 52d9923c-d046-424b-b4af-68dee322933b · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.589017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.776492Z digest=sha256:947372f61bc454c5799ccb019a7f2e31bf6d29733934bedc6add094802ae49a2

Observation e5a4718e-8a2d-4a26-be94-da876beb5e8e · outbound

This paper cites Consistency models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Consistency models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.391621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.839694Z digest=sha256:7df4e737b717321741bed65d1b1c755af5c245f22af6f36c464edbc8d8010f57

Observation b152212b-1d78-43b0-bb85-10d3a4ece214 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow straight and fast: Learning to generate and transfer data with rectified flow,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.190716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:15.926474Z digest=sha256:e5ba98a53dcfe96a2ae5c4e2d0c29be0018da64c741b0dc472bceaf510df2863

Observation f858a989-c3cc-448b-a3f4-7e1ede8df9f3 · outbound

This paper cites Comospeech: One-step speech and singing voice synthesis via consistency model,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Comospeech: One-step speech and singing voice synthesis via consistency model,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.931419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.004306Z digest=sha256:8b253427c906783e712a926ac6a83f866ce5dada8eb9236f31bbcee98182f2c8

Observation e52ae0af-4104-4462-99cb-d18419bc7542 · outbound

This paper cites Reflow- tts: A rectified flow model for high-fidelity text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Reflow- tts: A rectified flow model for high-fidelity text-to-speech,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.730383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.052933Z digest=sha256:eceb7810ab0daa09949d08f97256cc887a01dc34c2ae86750fec5f1683d3eb6a

Observation da6a2ad3-611e-4da0-9196-71b71c128ca1 · outbound

This paper cites V oiceflow: Efficient text- to-speech with rectified flow matching,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V oiceflow: Efficient text- to-speech with rectified flow matching,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.567457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.144964Z digest=sha256:f42940a13634e453e0d1441d715b35ba64ae9e03c22ab8a0d2a4ae598bb4589f

Observation 6e52ee7d-d6cf-49f8-a13f-9e2be46f25db · outbound

This paper cites Flashspeech: Efficient zero-shot speech synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flashspeech: Efficient zero-shot speech synthesis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.354922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.189257Z digest=sha256:fe6a96eef6733dfc2a1e8f6efe08f42663fbaa04cf683faa870934a6bb295087

Observation fd3bdbbb-af80-4d3d-beef-ca010a89d28b · outbound

This paper cites Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.136763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.251832Z digest=sha256:6196ab7c94a945e37a0f058d2da53a87e38f575f6508ea920a0a258b30f038b9

Observation 92523cae-b210-41cc-ae14-2d9e03326ba3 · outbound

This paper cites Lightspeech: Lightweight and fast text to speech with neural architecture search,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Lightspeech: Lightweight and fast text to speech with neural architecture search,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.922989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.341634Z digest=sha256:458c5b09efe3087cd9508ae06a0d010fc90e1b0e5aee87d97ebc1e2eb9f3f512

Observation ceb84728-7a54-48fa-8c99-e7d41459232b · outbound

This paper cites Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.724371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.410344Z digest=sha256:96c1825419018f1bc77fa8824a501a975a333296d50878b269bea67808a9ed63

Observation a08e7b59-6a11-42a9-b2db-21d3c47cefe3 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Common voice: A massively-multilingual speech corpus,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.561055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.499791Z digest=sha256:e9a025397f7bec6ebeeda946003c8abbd8a621810958a0cf09d488e0e9808e30

Observation 4bf2354c-ea56-44df-9c1f-b1f82f44ebda · outbound

This paper cites Didispeech: A large scale mandarin speech corpus,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Didispeech: A large scale mandarin speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.414015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:16.572166Z digest=sha256:420361802e379f821099f2d7d71b892c58ea4432e2c28ed61aaf5cd395c202ac

Observation b922f287-09de-418f-95d3-a394b95e554c · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.635875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.635875Z digest=sha256:136f75fdc41fe7cc64851ff6937a612472db420375240f0d6d0a3078ff7242e1

Observation 66d38499-0d1c-4dba-9295-38bca924f20c · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Robust speech recognition via large-scale weak supervision,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.715337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.715337Z digest=sha256:e1dce7096207682428975a3530b4e01b68a78ad8cb3861118b988e519ba3da90

Observation 28641d09-7533-4f30-9ff5-06211cc5fb0e · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.775147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.775147Z digest=sha256:7ee82d33eba325372a7539d9e78b00396a215f35f0362492c68f96e6f727256d

Observation 87dc4532-6b2a-4054-85d9-5caaf380e904 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.855648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.855648Z digest=sha256:6d44993587d824a6d1e650fee32ffb25adeb4633147bd127b36e0ec74da38ee4

Observation bf4eb6b6-3f38-412d-a71e-47c23e715ee4 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.920932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.920932Z digest=sha256:1e96542a72f5aff203f24f234c15a05c397ed717586249be9d677422ddeec5d4

Observation 523f1dfa-d805-481b-a534-f640a4ae254a · outbound

This paper cites Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.031709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.031709Z digest=sha256:3f0aee316f1a0385c185cfd6b0adec8ff191d49ce7fa2df74c51a97c2b64cbba

Observation 31e12026-c525-468c-a783-f51f3b05bb96 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.197912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:17.111078Z digest=sha256:5a48df538ee321c0d958608186919449912358a6401d849c9beaaeaf29269e7f

Observation 263e86b0-e18e-426c-8237-e6dcd3010d8e · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.180878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.180878Z digest=sha256:62e079c7c25d8709dd85dea6330f1e7910fd80b6b919c6b1e73512ec44dd0eda

Observation 80c5d825-271b-4e00-9722-bc71a0a9a635 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.249552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.249552Z digest=sha256:f8776ff8a8455822520f024e77b9ba288ca60ba3b6020b81282de1096beea87a

Observation 45102f47-a413-4735-b698-7e1c2858eaa9 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.314055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.314055Z digest=sha256:f87f0e1d23bdd16a1f83134fa6c3a34b13cdae5fabd3f2e5b06b52065b393a80

Observation 11b0aff7-79ae-4cc5-9a05-3e065b821938 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.058741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:17.387188Z digest=sha256:d797954ed08a2298ef2fb90c8a37a14b656df3938822e2fefd71dee8db1390b9

Observation df6b7d1e-4c64-4752-ab52-b75bfd80e94d · outbound

This paper cites Amphion: an open-source audio, music, and speech generation toolkit,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Amphion: an open-source audio, music, and speech generation toolkit,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:17.878084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:45:17.467537Z digest=sha256:2c2379cb58ea00fbc6b296d42fe3fed78f06669d3585352094d31a1aa9114cb8

Pith citing papers

Observation 1f256cb3-22a5-4e2f-be92-73228fa1566d · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.673585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:ef5d7bff524cd7a72a76d26f1bed2c9e0af14cab41d758d03b7aa522308266ec

Observation ab2e20d5-7d2d-4ad0-b454-c474d228d099 · inbound

Universal Speech Content Factorization cites this paper.

Universal Speech Content Factorization ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T12:21:48.698333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:21:48.698333Z digest=sha256:238b4ea26193fe60ccda1db66f194f1624be03ba6301440ebbd1fd64c675f15e

Observation ef4f0b28-14ef-48b4-aad7-45836011a635 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.708672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:564031dbae870f29a09c813e2ac456c09a68779efeb0f9221f5099e97dded618

Observation a7b0fc38-7d06-4288-9379-6e9f4c2ce3d3 · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.048637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:1c726d9f61a51394964aa0a735941a12cee8f0a0b3ec6bd3e6811d09f3be264c

Observation 93f45acc-75fe-4625-954f-7155ac151c43 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:36:45.127244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:fac00118e3120b7b1a08c6f80b86e3283aac808cf37989f9455cbdd731e0e932

Observation a148795a-e901-425f-921a-39a80ed99b83 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.821926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:3b46dd310748975c09f3fe8fe3d92fbb8c71c6a5fc338e75d674625a03e6759e

Observation fce86089-99f4-4aab-9876-524fa05de4cd · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.049565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:e980e13d24b301aee14e9bf770294a239f427ba98d83b1c3f171ff6dfb4eedd9

Observation 456e7331-2449-4af2-a6c5-fa6a9ce4ec35 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.093666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:c89876f2375c1f7b74d62c5dbe6f0ce9b17a3825ac220cee90f7f99f4d0c9231

Observation 35be436f-b14d-462d-9f24-c9059434f62d · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.787929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:9b5375b811af1b1f1a98f9eb2dc961c833507800461b67f7bd5f1560086a9d09

Observation d30edfa3-c12d-4673-9a20-d17c3e8d3216 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:47:19.794924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:c6a1510c8710d5881740d2111815833ca07fed0da48852c3af15e77d0f36f98f

Observation d1c78e43-7078-41ac-acf0-683a37a18d96 · inbound

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation cites this paper.

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:19.602737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:14:08.243894Z digest=sha256:9e39c381079edcca09ccaf6cad1aa255f3b872920f0a5a5addf331eee31089c9

Observation a83fb087-1b43-4394-b06b-ed907f6a3bd5 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.898147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:ab31cb2116987efd65357f91ab5fce87dec0ecbf46f4455c4076422af77e1bcd

Observation 7a947619-156c-491b-a449-a4279363051c · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:38dfdbe1f0ef700317e7475d610c5a68fc69691499c7fd61f22018fc356e4bf3

Observation 0297f999-eccb-423d-876a-5296308efd40 · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T02:22:47.820537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:22:47.820537Z digest=sha256:39091cb1f75eb84dab41a30599dc8edc6fd057ea9beb7ef9d659363578ff3a30

Observation e07bb246-f5cc-41a3-9d83-cc58e8eec40d · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T07:39:23.947687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:39:23.947687Z digest=sha256:31965a1f1318e80183c95291c5229319ec2373cf0a12384377c7425813779e6f

Observation 51c0f031-9d9e-4949-9112-f3042e5cf35e · inbound

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model cites this paper.

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T22:26:14.049693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:26:14.049693Z digest=sha256:ff4edb19acaab747502e554a75652bdea131c7338c14a5598bb4b9d8894b52ed

Observation 8cdbe7ad-c11e-466b-814e-1739038e6f0b · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:41.967244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:41.967244Z digest=sha256:72937b6ca6d2e09b3a6cfa0bfa258ed2276af0cab813096ed4b393d12b5f7d6d

Observation d48ff219-ef5f-4097-b213-fb27449d39b1 · inbound

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents cites this paper.

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:47.834426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:47.834426Z digest=sha256:97051774b126558bf6f78d0b85d8a2a184bdb4667e01a068c058184a43790a7d

Observation e9b42b5c-8d5f-42b7-b6cc-28bd80a5fdf0 · inbound

Luna-TTS Family Technical Report cites this paper.

Luna-TTS Family Technical Report ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.263064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.263064Z digest=sha256:e535431decfd19f39ea5ddbb543366618eaed73687de6d727a8eeee015147fd8