Pith. sign in

Paper Citation Record · LEDGER

HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2410.04380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.04380 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:26:56.531216Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T19:02:43.552535Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7e9dd3b9-8940-4494-a023-8fe45fcc6085 · inbound

Emotional Face-to-Speech cites this paper.

Emotional Face-to-Speech HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.825414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.825414Z digest=sha256:1e824bfecf5329359022577c2128216261b982449132c60507b086319cdf04fe

Observation ef5bc3fb-070f-4386-9efb-2906ed5fbf84 · inbound

Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding cites this paper.

Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:26:56.531216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:26:56.531216Z digest=sha256:1159181d52b1460ae3fba11e0edc5663d2d685db7837de730814d11b534e4b05

Observation 992085bb-1e13-4e77-a897-c0c0e2e2511d · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.181491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.181491Z digest=sha256:b249a4ce9441b37e54bcfb32967c3bdaabf1aa5e0c44c696397c596c5bd1f916

Observation 7f4a5b07-c22b-41c6-bf22-5d42a07dcd83 · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.424043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.424043Z digest=sha256:631560d48e0f4bd27de413e83d3324fa73fe72a41d1fe73ecf9eefa05a17bc1d

Observation ca2f0a00-cd01-453c-ab18-6e562a1fbf90 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.722764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.722764Z digest=sha256:3f9ed0d5788a441eb6e56eeea9dced399f48d0e53ac12f529c28077666f42ba1

Observation 2347a4f6-b275-49d8-9020-0220823e8fcd · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.108916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:8f36bfd645cd4ce459bf3ebc650c792603d53b4ed1f3d4ee33ac340d49c5bbcb

Observation 148aecbb-bd2a-4eb1-b797-1cb8b87f8337 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:997d6c1bb4e6bf4ce8e70bb64eaf8b2aea962f5e3388013aa65296a583a92c74

Observation cb1e853e-ccfe-4e78-a613-bd035821d010 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.554199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:636fa49338dccd693b5269c35cddfe64ea013cf7c776e09dc326a45073bfbe30