Pith. sign in

Paper Citation Record · LEDGER

HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2311.12454.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.12454 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.254583Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T12:26:37.491138Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1b46cae5-38b9-4e85-b7a8-66bdc50c31c2 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.493296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:ade9d98cebcd09b7bba9a677c0c335cae59837b58c72eacf9e49fd08155ec3df

Observation 49e1554d-4460-4704-a07a-acb5915ffb54 · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.254583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.254583Z digest=sha256:a28ee8cec1d13edbf4188c819dd6f8963237537bb8f48575d37c4a92642f119a

Observation 3cbbd4e7-709a-402c-b2e3-a61fcb7b3d06 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:55.976737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:55.976737Z digest=sha256:d5fc45cae6a3df172e02212cbc01f4a10fa2cd153b379bcfc90dc52522bf6ec0

Observation 465a0c8c-bb1e-4e75-8021-eb5edcc90a2c · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.325815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.325815Z digest=sha256:ee3400b2f44bc9c0f58a241217d2fa321d5b23cca0b2ed3404e0b51b5d77cb73

Observation 361f327e-a2bb-42b2-b314-ec4cf2d7f6ff · inbound

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment cites this paper.

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:07.436262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:14:07.436262Z digest=sha256:f060917778b1f6de7ae409f46aed1fd8ddb5382668e888336deb18f893a00c5b

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:33.784245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:33.784245Z digest=sha256:ba6bcf5a5877ffed3e74af090489d0369035e162ee7c16bb26fa7db0c8d36bd0