Pith. sign in

Paper Citation Record · LEDGER

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2306.07691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.07691 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:48:38.355724Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T22:46:12.436746Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2ed63fa2-7dcd-4bf0-9e5b-ff260e294d0b · inbound

High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR cites this paper.

High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:48:38.355724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:48:38.355724Z digest=sha256:f6814f2c0bd8a88bec643d64076204cd8d1143d2723b0d00e597fa2aa3038682

Observation dc31d27d-02e9-4448-a840-678c6b080808 · inbound

VocalCrypt: Novel Active Defense Against Deepfake Voice Based on Masking Effect cites this paper.

VocalCrypt: Novel Active Defense Against Deepfake Voice Based on Masking Effect StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T18:34:13.154982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:34:13.154982Z digest=sha256:c21a380817de1d326ab5c8f2d4103d5dbdb71d7e518ab379e24ede01d33d10b5

Observation ffe792b9-1823-40a8-b57d-84b5f44a1e5a · inbound

Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation cites this paper.

Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:56.300139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:56.300139Z digest=sha256:cf5a4063ad109f5cbcc268a5141ebcda0ef73a845dc4e9562bb22968ef93ee5e

Observation b54c6fbb-5bc7-4da3-bc43-931e77ecb537 · inbound

Spoken question answering for visual queries cites this paper.

Spoken question answering for visual queries StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.339674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.339674Z digest=sha256:c27770b68cee90b878f39424b102908fb81a012c42615ce518dba7e0a01e4553

Observation 87337e4f-1806-45e6-bfb4-bede7ddf54e3 · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.153909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.153909Z digest=sha256:487eddcf4eee62265a59193ba8f0aa248b703c68518231adafcd7fd551c4d0ca

Observation e9a9e577-e299-4b9a-a7d3-db4adf7fd36a · inbound

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative cites this paper.

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:46:12.440333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T02:27:05.776290Z digest=sha256:b84f8da5e9bbdbd644bd87b32dd45ef9010dd7531068ad03959fd57a4ddd4ae8