Pith. sign in

Paper Citation Record · LEDGER

HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2311.12454.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.12454 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:32:24.075261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T12:26:37.491138Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1b46cae5-38b9-4e85-b7a8-66bdc50c31c2 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.493296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:600482366ed06b777e195966b334c6cc28f2bde2b26d2e5c3f07468ff293a717

Observation ecc71c65-57fa-416a-ba47-2d1417464cdc · inbound

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey cites this paper.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.075261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.075261Z digest=sha256:d70018454df0f49f5a11137968f46fc5f5d27bdbe6c37b6089d864ff489646a8

Observation 06087c3f-9f4c-4b7e-a902-b7952bf232b7 · inbound

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron cites this paper.

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:15.815810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:15.815810Z digest=sha256:620b29b6c0431e0652864009d94c3341b3d9fe9d27abb1a347052fe0f0cb213f

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.827008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.827008Z digest=sha256:a4a78e7c3e425e1825c975c95880120a4a613b823244f2cad9cb000862d683d6

Observation ceaace19-f2da-4f7a-97fc-c90eb85c1b64 · inbound

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis cites this paper.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.772217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.772217Z digest=sha256:dea2a7077cc95fc130709d36696bc2e326a619630d3c1012bb46dbd9b3b59bcb

Observation 1e1a0fcb-be09-4d73-934a-3e235b15ca80 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.603332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.603332Z digest=sha256:80620d6f71ed042e8211396bce982dcc42501ceb5950244b5e51da4d6e1c0578

Observation 49e1554d-4460-4704-a07a-acb5915ffb54 · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.254583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.254583Z digest=sha256:a69e2df17071b3294dd75876ce7473158a67dfa3441280fd8ebf5961b11d67d2

Observation 3cbbd4e7-709a-402c-b2e3-a61fcb7b3d06 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:55.976737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:55.976737Z digest=sha256:390df95ed8d2aa1d2865e808981709cca25678854bac948c0ab5ae20424afad0

Observation 465a0c8c-bb1e-4e75-8021-eb5edcc90a2c · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.325815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.325815Z digest=sha256:bb8699409fda5265b0ec40e12c9846b091c18439a0b32dff06e5875a613569bf

Observation 361f327e-a2bb-42b2-b314-ec4cf2d7f6ff · inbound

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment cites this paper.

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:07.436262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:14:07.436262Z digest=sha256:f96dada594ac3b62387fddf297ff9052c95af6130f53b15791e75aee5db82794

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:33.784245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:33.784245Z digest=sha256:e7530077fd9c9604587411016d6435168ff19a5a236bbc222adaf0f1b090cf28