Pith. sign in

Paper Citation Record · LEDGER

Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2106.06103.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.06103 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:29:39.093644Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:29:51.975259Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ed9707c2-50d0-427e-960a-34d51717a85c · inbound

EmoNews: A Spoken Dialogue System for Expressive News Conversations cites this paper.

EmoNews: A Spoken Dialogue System for Expressive News Conversations Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:29:39.093644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:29:39.093644Z digest=sha256:cbee587d61e69a0dc6e3f3aa55403d970dfec0fef123bf6e0471e27b56552655

Observation 4b0dbbc5-4b3f-4261-bc0a-71ce1117497b · inbound

Conversations with Andrea: Visitors' Opinions on Android Robots in a Museum cites this paper.

Conversations with Andrea: Visitors' Opinions on Android Robots in a Museum Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:16.986708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:16.986708Z digest=sha256:aa839e3bfc40ef82bd14b1b4f783e439bddcd39cb95fe34b905b265f9c70c913

Observation d5ce8eca-58e1-46c1-92ab-e59d5d3f4173 · inbound

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages cites this paper.

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:29.789462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:14:29.789462Z digest=sha256:9d47574d26b2a0915567f8d0c90d9be61d9b1735e9d1f942bb7573bef77c0ce1

Observation 867aa1ca-1fb9-4924-9e32-ac2e1e0b9e10 · inbound

Marco-Voice Technical Report cites this paper.

Marco-Voice Technical Report Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:05.972642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:05.972642Z digest=sha256:71e88c8e22be12a3566379fb22541ece2d1afdb861d34ff556e74f5d41be044b

Observation 862280ad-d7fe-4f7c-8889-4dc561fe5c4f · inbound

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features cites this paper.

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:17:25.427389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:17:25.427389Z digest=sha256:4c194ce32cfb56b29e4b73edb3d981ccdd3740a2e08dda9be5fe35af7093532f

Observation b4cb3d38-8162-4668-8c14-50cc29b5a089 · inbound

IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA) cites this paper.

IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA) Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:18:30.029883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T00:15:10.797922Z digest=sha256:ca2348782166ebcfd78f8d4c4ab799f5730f128927d13c9cfaa92c73144448a2

Observation 47a31dd2-a8d2-47bc-b0d9-d24c08060bfc · inbound

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection cites this paper.

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:29:51.977508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T06:45:08.471913Z digest=sha256:07fa813ff9b83c65c3ee64ab18a3b48e3c238e7d10c41519426008e9a0959121

Observation 2a8d74f5-5cfa-43ef-aa50-2581398d60bc · inbound

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings cites this paper.

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:22.075952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T01:21:02.597103Z digest=sha256:d7e4892c23fe749ff0001bf0a29ec082a6520dddaeac1406d7d06efb5318f40b

Observation 0cb4d44e-5b4c-48c5-b840-62f658273daa · inbound

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language cites this paper.

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T18:11:36.626189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:11:36.626189Z digest=sha256:3213ae7b93595883e1462c0d5cf084afad18c45319a955b4d3454033ae9313fe