Pith. sign in

Paper Citation Record · LEDGER

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 17 inbound Pith citation observations for arXiv:2506.13053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13053 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:45:17.467537Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:20:41.967244Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:34.896349Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46dd6f12-388d-47f2-b4d5-1b58801c4773 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Neural codec language models are zero-shot text to speech synthesizers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:24.192871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:13.214738Z digest=sha256:d2dc2a7ced9adee084296e29bfe0a14cad3448ae1d2671980464bf5cb475a771

Observation 2cbfe618-9751-4ea9-90ad-f4aac7513486 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.288815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.288815Z digest=sha256:e44ec466d7ea16e4d461b9f0ff012efc262094269b964d1d81a71e59b6911cb1

Observation d4840c43-7c3b-4257-980b-e308b9d53508 · outbound

This paper cites E2 tts: Embarrassingly easy fully non- autoregressive zero-shot tts,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching E2 tts: Embarrassingly easy fully non- autoregressive zero-shot tts,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:24.038565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:13.380779Z digest=sha256:a432604c4f3cbb669938443ba407958af37a763f106d1ef9fe80c99d61687dae

Observation 19e19724-5546-4746-8948-d7be6e8427d4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.443312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.443312Z digest=sha256:18e3642a0e02c14f285c363f12f9d97c35d11f13c2963461048fe70bd769ccb9

Observation c4ee4589-42be-40e8-b106-1220f2f921ad · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.893891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:13.540878Z digest=sha256:1cbf07686b3d19078bf1077109503c33346fe4fcea5d52f9aefed1f6858a448f

Observation 47614f9a-f3b0-4357-b723-32f725028c9f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.608161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.608161Z digest=sha256:36c6d6f1d2afc4604561390217c650e0076286c198395e56a85cb4dc5eb6db08

Observation 811b01f7-0773-4610-a07c-7bf31ea74375 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.675822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.675822Z digest=sha256:71d57640de238220299d537b378f02fa61b3fb749da68f48686ec2e9dae18f8d

Observation 48f1086b-b0ef-4d23-aeb1-644adc8484d4 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Libritts: A corpus derived from librispeech for text-to-speech,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.776702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:13.755949Z digest=sha256:7bbc209d2caeffc83dbebb64209ae4f1d4395364aa79bcdf9123c17d9664c3e4

Observation 1a7c5959-0e17-4458-b945-020459b91c96 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.847783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.847783Z digest=sha256:6d2724e33325273ec32f62022e714863f88cc871fcd0a2e6139e9dcb264f845e

Observation e190dbb3-1fbf-4a1f-9e04-e730369071e3 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.630165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:13.919856Z digest=sha256:af5a6159abd549534e1cb9c742e2e88dfcfe8cd0b5c3fc45615e543628a4ae3f

Observation c4b312cb-f7a2-492f-b2c2-3ea22dbe1a6a · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero- shot speech and singing synthesizers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Naturalspeech 2: Latent diffusion models are natural and zero- shot speech and singing synthesizers,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.479251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:13.988814Z digest=sha256:4c5577211bbc315ffac2b66bc24d62ab68a25a50698286d4a186a98d374a4d37

Observation 0648c2e9-ec80-40e7-b40c-854a16fed321 · outbound

This paper cites Flow matching for generative modeling,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow matching for generative modeling,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.333080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:14.064823Z digest=sha256:9a5b56517f02093ddacd4a67f7ffcb5f3f1fe8d15c305e1c7375b1a280f38579

Observation c726225b-2ddf-465f-9cf5-1cc4fdbf401a · outbound

This paper cites Sf- speech: Straightened flow for zero-shot voice clone,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Sf- speech: Straightened flow for zero-shot voice clone,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.200164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:14.143790Z digest=sha256:0fb309496df211ccb5a068ee44dfb44edd636644892aceb9dc110fb8a66c1fbe

Observation f7b3f264-2b8f-494d-a920-a01aff7949ae · outbound

This paper cites P-flow: A fast and data-efficient zero-shot tts through speech prompting,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching P-flow: A fast and data-efficient zero-shot tts through speech prompting,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:23.080631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:14.242517Z digest=sha256:4a73cb942520d104dcb36a735065e11003fa73fcb02fe736e750f9809ea64497

Observation 8e3f0334-850f-41da-9b83-dbe729364cf0 · outbound

This paper cites Attention is all you need,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Attention is all you need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.325866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.325866Z digest=sha256:de58f556bfc6c43ffcb888e358567bf96c3f785d60c1fb1cd0a99bae8e12eee4

Observation 1e0b8f32-b682-4c29-ba91-3741fc249547 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Zipformer: A faster and better encoder for automatic speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.897794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:14.417592Z digest=sha256:8caea256fbea2323a157b7babeaea4e29fc836fc8316d6519da721faa5a680ab

Observation 2110306a-b124-47bc-be6c-8bd87d96343e · outbound

This paper cites Classifier-free diffusion guidance,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Classifier-free diffusion guidance,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.760185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:14.504825Z digest=sha256:361f665fb19f84c3800a283b120e2761e3d564c914f95721eabbc2dfbdc1c65d

Observation 170f12b1-dcbb-4192-bdf7-bbe89d575b1c · outbound

This paper cites Freeu: Free lunch in diffusion u-net,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Freeu: Free lunch in diffusion u-net,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.605205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:14.572199Z digest=sha256:cbe8e031e058714f2ace7e118577527036b7ccb3e3301ae57f91f4b8fb27dcd4

Observation 4a2bb12a-be56-4f7b-9cf2-5cc4415addbf · outbound

This paper cites U-dits: Downsample tokens in u-shaped diffusion transformers,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching U-dits: Downsample tokens in u-shaped diffusion transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.210548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:14.658013Z digest=sha256:71819dda3788c123129035218180cac0c279fa1e2a4c4fe680386ba5b5b18b1f

Observation 1fad3369-6cbe-4ed7-be2d-5ae2ffe0c242 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Fastspeech: Fast, robust and controllable text to speech,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.731869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.731869Z digest=sha256:ba4715c5e22b33ff70529e4f53e46bf775b679017cab2035c7bb75a1924a243d

Observation 94ec5498-4dba-42f3-ae84-7ae06ad8cc7c · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Conformer: Convolution-augmented transformer for speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:22.047617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:14.832424Z digest=sha256:1009f0dcd1d0ca5b7da940c33de46d40b68edf56b3a62e7e4f916dbde6a5a6d8

Observation b9194bd8-2d98-4212-8bfe-fd2d1567bdf4 · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:14.930725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:14.930725Z digest=sha256:a045de6c1c4fc7915e837ebf5403ccf764366e737e8ea3e99b600ace51e2e0b1

Observation 64bcbf28-fe5b-4eff-a317-75f4b2b627ca · outbound

This paper cites Flow-tts: A non-autoregressive network for text to speech based on flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow-tts: A non-autoregressive network for text to speech based on flow,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.877248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.033007Z digest=sha256:ac261934a22dac0e342b051a0f3256dedaa086b8b827ce5afb64c4ce9d62fc15

Observation a84aa383-803f-4188-93d9-7e3c6359d2ab · outbound

This paper cites Simple- speech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Simple- speech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.707431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.151321Z digest=sha256:94dbd6a78bbe487fb99c017e3ccc3a8cb8f22b33d9625236cb29d599b7d51f32

Observation e5c11d3b-3dc1-4568-b894-f926fe03ec62 · outbound

This paper cites DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.441890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.249852Z digest=sha256:5c92e1f26116c8e575d3e46dba93bf7c27bccb83071a3be7f9e9a17c4a0e9113

Observation 9954c653-8243-48f0-8f7c-bee76df3b892 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Convnext v2: Co-designing and scaling convnets with masked autoencoders,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:15.372934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:15.372934Z digest=sha256:e16ddaa4de33935ad55b8ed192281b3e668964e34bdbeb2d1b5df782d8efaf8b

Observation 38ae1325-a8b2-453f-a1ba-4bc4e81cc555 · outbound

This paper cites On distillation of guided diffusion models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching On distillation of guided diffusion models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.229679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.438075Z digest=sha256:281588b0ac868c51590e96bee5811cbd49e9b5eb578cc9e60c356591ffd43031

Observation 8a9aeb93-b8ac-45a2-9e03-00166a481c25 · outbound

This paper cites Tacotron: Towards end-to- end speech synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Tacotron: Towards end-to- end speech synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:21.032849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.495205Z digest=sha256:06ccd2d8a1fa31a45e8609e9d61aca04a135ed422d3549d4b793ded31fb9a3b7

Observation c94e5b87-c82e-415e-b725-0232d8a39a93 · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:15.573777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:15.573777Z digest=sha256:2fcd967ea58e92573df1bbd20fa93fdd3e5c96bf990c045f20f6a311c922fd73

Observation 40a0df60-3838-450b-ae44-fdff04008690 · outbound

This paper cites Revisiting Over-Smoothness in Text to Speech.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Revisiting Over-Smoothness in Text to Speech

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:45:17.675731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.639463Z digest=sha256:4ace7fbb51c5ee25f760305bc10503bb74b9e587a770d695841537bf62171c15

Observation e61555de-6d6c-4486-9be1-8c877b4ac97c · outbound

This paper cites Grad- tts: A diffusion probabilistic model for text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Grad- tts: A diffusion probabilistic model for text-to-speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.775881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.700325Z digest=sha256:ddd7f24ab23badc82c107d83c8f146c4aea7ead17e51e0c004f4dabecbc04107

Observation 52d9923c-d046-424b-b4af-68dee322933b · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.589017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.776492Z digest=sha256:9f0fe6555108029bfda37515f6e03b07452e7348781844ab5985e10328007a5a

Observation e5a4718e-8a2d-4a26-be94-da876beb5e8e · outbound

This paper cites Consistency models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Consistency models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.391621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.839694Z digest=sha256:d66fd6d010b298fb038374a66ab3158cd0b337344b96082c248f7ef77194a574

Observation b152212b-1d78-43b0-bb85-10d3a4ece214 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flow straight and fast: Learning to generate and transfer data with rectified flow,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:20.190716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:15.926474Z digest=sha256:ec10ff76c904115bbc61dd3cca2e38769918dc06e2429e0b3e23ff42d61018ca

Observation f858a989-c3cc-448b-a3f4-7e1ede8df9f3 · outbound

This paper cites Comospeech: One-step speech and singing voice synthesis via consistency model,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Comospeech: One-step speech and singing voice synthesis via consistency model,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.931419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.004306Z digest=sha256:bee48f73e49995d32c0a9d691372ab8e794c6416bd6a3c248ea1da946a37f6b8

Observation e52ae0af-4104-4462-99cb-d18419bc7542 · outbound

This paper cites Reflow- tts: A rectified flow model for high-fidelity text-to-speech,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Reflow- tts: A rectified flow model for high-fidelity text-to-speech,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.730383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.052933Z digest=sha256:ef6cf8f0ae1e4b44b5c8c37e6a0ca2b8f143d50a4f628be9560b40ea5e67f64d

Observation da6a2ad3-611e-4da0-9196-71b71c128ca1 · outbound

This paper cites V oiceflow: Efficient text- to-speech with rectified flow matching,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V oiceflow: Efficient text- to-speech with rectified flow matching,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.567457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.144964Z digest=sha256:908f1e9deb90c38e19d564f523977b7ee2fe23b985f835ae655c6ae6dbd2ba75

Observation 6e52ee7d-d6cf-49f8-a13f-9e2be46f25db · outbound

This paper cites Flashspeech: Efficient zero-shot speech synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Flashspeech: Efficient zero-shot speech synthesis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.354922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.189257Z digest=sha256:913db8a7426a03c130c35bc2d797f1da917e7aab65ac3d05d1d35e87863d2591

Observation fd3bdbbb-af80-4d3d-beef-ca010a89d28b · outbound

This paper cites Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:19.136763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.251832Z digest=sha256:204af56da3fc401c4e093cebe30b080d4b1f99ebcce72294bcfa72001e0dbcd0

Observation 92523cae-b210-41cc-ae14-2d9e03326ba3 · outbound

This paper cites Lightspeech: Lightweight and fast text to speech with neural architecture search,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Lightspeech: Lightweight and fast text to speech with neural architecture search,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.922989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.341634Z digest=sha256:be4e02983859cd5590f372ae75e79fb770c6dc0b437443e67f1f5cd3fc95c5a5

Observation ceb84728-7a54-48fa-8c99-e7d41459232b · outbound

This paper cites Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.724371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.410344Z digest=sha256:31e832b09841603bacbc201999c3ba3ec303f462dccae579cb1cf16110c4374e

Observation a08e7b59-6a11-42a9-b2db-21d3c47cefe3 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Common voice: A massively-multilingual speech corpus,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.561055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.499791Z digest=sha256:6bd31e5786a3067ce1254dd09067a30af074f2bfc5f2c9dd8867aad8e889dba9

Observation 4bf2354c-ea56-44df-9c1f-b1f82f44ebda · outbound

This paper cites Didispeech: A large scale mandarin speech corpus,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Didispeech: A large scale mandarin speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.414015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:16.572166Z digest=sha256:fc87dfecedadd6260fd6c4564b71e53c83e9f97a1589c29850ae1b65dbca1414

Observation b922f287-09de-418f-95d3-a394b95e554c · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.635875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.635875Z digest=sha256:226d6749577069776d916adec027945b2ee5bc181dee5388c7b93cf93dfe27ce

Observation 66d38499-0d1c-4dba-9295-38bca924f20c · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Robust speech recognition via large-scale weak supervision,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.715337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.715337Z digest=sha256:1f7c8c103fdf7b4b9ff5c1e7df0442ea0103262d043676920acdd24a1b52605d

Observation 28641d09-7533-4f30-9ff5-06211cc5fb0e · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.775147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.775147Z digest=sha256:e5df0ead2550e6dcff36de047790230c1d556a44730e594a4de3fe89ae6e7e22

Observation 87dc4532-6b2a-4054-85d9-5caaf380e904 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.855648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.855648Z digest=sha256:179588aa873df21b91dd8e93de3d9ba70a4672b498d5f687dd5783e539f037bc

Observation bf4eb6b6-3f38-412d-a71e-47c23e715ee4 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:16.920932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:16.920932Z digest=sha256:bfc4cd01e0546ab3c3e50bd1f547a2ae5b17df06fee26a88bf95ba4b3aa28cf8

Observation 523f1dfa-d805-481b-a534-f640a4ae254a · outbound

This paper cites Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.031709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.031709Z digest=sha256:a3bda89fb78d3ae9d7d1e01881a1fc0089033f079f67b90807351beaeb4dbe0c

Observation 31e12026-c525-468c-a783-f51f3b05bb96 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.197912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:17.111078Z digest=sha256:02702fdf598a4916f20cf8bada145c480cf2b60051728090553eb668b601f5ab

Observation 263e86b0-e18e-426c-8237-e6dcd3010d8e · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.180878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.180878Z digest=sha256:038c63df969fa2ac02b8cdf8347200d88209ff82d92dc4fb238a9cf705d93778

Observation 80c5d825-271b-4e00-9722-bc71a0a9a635 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.249552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.249552Z digest=sha256:1f525218909be144a2712875d7e7aba1350124c69faa51c1b9b813878b34c212

Observation 45102f47-a413-4735-b698-7e1c2858eaa9 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:17.314055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:17.314055Z digest=sha256:a4872e6a10cc23d4f4f8f54b526892f5f3fca308f63ef96a9c93ea5f1e5ab00f

Observation 11b0aff7-79ae-4cc5-9a05-3e065b821938 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:18.058741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:17.387188Z digest=sha256:d8f991de398e585fbff8bd1b66a9a767aee95f74cfb604468de62ed656bde82f

Observation df6b7d1e-4c64-4752-ab52-b75bfd80e94d · outbound

This paper cites Amphion: an open-source audio, music, and speech generation toolkit,.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Amphion: an open-source audio, music, and speech generation toolkit,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:45:17.878084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:45:17.467537Z digest=sha256:5a85c2bd8b2e90ac2edb51ed598fa0cca1134a7eebaeba508e7ab07f819b8d70

Pith citing papers

Observation 1f256cb3-22a5-4e2f-be92-73228fa1566d · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.673585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:87a91ae083dc8105408446dfe48e81de719c9a332eb2291790d692c828087471

Observation ab2e20d5-7d2d-4ad0-b454-c474d228d099 · inbound

Universal Speech Content Factorization cites this paper.

Universal Speech Content Factorization ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T12:21:48.698333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:21:48.698333Z digest=sha256:6ce8e74e30804ed8ca29a21e08ce1281395273297d8f7cc15bd70efa33e6720a

Observation ef4f0b28-14ef-48b4-aad7-45836011a635 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.708672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:ad73743257917fec92698401f5269d6e842a5f3e1ce995ef1b5566cae5eddc6a

Observation a7b0fc38-7d06-4288-9379-6e9f4c2ce3d3 · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.048637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:e845bc23c3d3c61f4379e50d844179562f32e704afc4d0861b530a5817f8f88e

Observation 93f45acc-75fe-4625-954f-7155ac151c43 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:36:45.127244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:37d5d98c84f0a8db074d7c6b14e3585076bed8cddf58cd138e2102aabeacbaae

Observation a148795a-e901-425f-921a-39a80ed99b83 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.821926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:fd06e00422ec3228a9ee3832afb4a0a9933385d6b1c2a45f95b25a07392736a0

Observation fce86089-99f4-4aab-9876-524fa05de4cd · inbound

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation cites this paper.

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.049565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T19:25:18.488377Z digest=sha256:0d4ab10e1cd7cca85db17ddcb1f06eb89ee2b82962c82ef435c98f80e3293fca

Observation 456e7331-2449-4af2-a6c5-fa6a9ce4ec35 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.093666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:c8aa18ea4256b5674598031d91dce2a9249f014cb5be974b4b5ae8f4d67ccf89

Observation 35be436f-b14d-462d-9f24-c9059434f62d · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.787929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:01cef772d67d41c9bfae4a8bb9bdca8b2c0516be39d03093ad1027fd00ef1e91

Observation d30edfa3-c12d-4673-9a20-d17c3e8d3216 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:47:19.794924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:e777a53137dcface242abf9c5c5c7f792a46a4cf032b1117605b3db983f6a907

Observation d1c78e43-7078-41ac-acf0-683a37a18d96 · inbound

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation cites this paper.

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:19.602737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:14:08.243894Z digest=sha256:0be55db3a282172927810ca5a26be4acbd041b3b75dae218785cf1968c48547d

Observation a83fb087-1b43-4394-b06b-ed907f6a3bd5 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.898147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:a2b01eae40f39448b2e05b1441fff69ef308e3a536d7d25fb17360bd5f1fd0ef

Observation 7a947619-156c-491b-a449-a4279363051c · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:402c7bbcca3882c29f2566dc486a41cc49311a233ae79e3d8857b6d21d03b5f2

Observation 0297f999-eccb-423d-876a-5296308efd40 · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T02:22:47.820537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:22:47.820537Z digest=sha256:6eff157f0e5064070f261444911bc5f6325a1f9235d645a9f43f190700f0e366

Observation e07bb246-f5cc-41a3-9d83-cc58e8eec40d · inbound

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis cites this paper.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T07:39:23.947687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:39:23.947687Z digest=sha256:3b6375136aa5836e3beb8c36ea836e664749f573e3b032e831cf5a7382012ed6

Observation 51c0f031-9d9e-4949-9112-f3042e5cf35e · inbound

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model cites this paper.

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T22:26:14.049693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:26:14.049693Z digest=sha256:db5de197f8446aa9c39975ccacc643b563f128b154642c93426b88d6dd9c995d

Observation 8cdbe7ad-c11e-466b-814e-1739038e6f0b · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:41.967244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:41.967244Z digest=sha256:78a042494c303f86fc5c3ab2f102a3fe3196c64d3069af5c6c06ca9044804549