Pith. sign in

Paper Citation Record · LEDGER

HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2410.04380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.04380 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:50:43.825414Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T19:02:43.552535Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7e9dd3b9-8940-4494-a023-8fe45fcc6085 · inbound

Emotional Face-to-Speech cites this paper.

Emotional Face-to-Speech HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.825414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.825414Z digest=sha256:4212f40f20ce877a06bf56c29a98025c2ad695cb5274babaed2099d5ed547128

Observation 992085bb-1e13-4e77-a897-c0c0e2e2511d · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.181491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.181491Z digest=sha256:d509ff8bc90093745af2553943f75bbb0b9ec2ea6ad5b7e77d642e72a48783b3

Observation 7f4a5b07-c22b-41c6-bf22-5d42a07dcd83 · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.424043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.424043Z digest=sha256:097af0a8a42f3c8d812e9efe67c097dbffc4e4311b9e47dfbb7fb74add7556f1

Observation ca2f0a00-cd01-453c-ab18-6e562a1fbf90 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.722764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.722764Z digest=sha256:996983b277adeeb35d5cd409f0010e8520a700cd965045c802337eceb422483b

Observation 2347a4f6-b275-49d8-9020-0220823e8fcd · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.108916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:668e2ff128e6fb981c2f50083bcca10de02a8e677d1ddffa4bbbaff1f283cae8

Observation 148aecbb-bd2a-4eb1-b797-1cb8b87f8337 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:38f06d6acf7473c0ffce79db96cbf2b3efe7038de9684bcfdfbc253759086de8

Observation cb1e853e-ccfe-4e78-a613-bd035821d010 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.554199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:b9b255e031701f14deb7618185cd06560bb985d1e7d2bbae129a6a7d5d8be31d