Pith. sign in

Paper Citation Record · LEDGER

ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2401.07333.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.07333 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:30:50.848560Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T06:06:41.523281Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d0f75a98-3c8f-4f61-8936-a580fe1f5d83 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.525280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:e4bfd2674fe1276cece612b9330870d919f8df358061845870c88c84a2112e8a

Observation 52a4f702-c67b-4136-bb99-3b2fed255e00 · inbound

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 cites this paper.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.337892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.337892Z digest=sha256:64ce44a6463db61fa86b338403a0a69db2ad39b0154d9a56dd600f31ebdf0b56

Observation b0ae0045-a581-40fc-a133-ecb6925a0a34 · inbound

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey cites this paper.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.101443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.101443Z digest=sha256:d16924863cd981e49659039f93873cfaaf3640934ea92711f75f8b7288c77c9a

Observation 51354e05-5798-4e18-8382-4f2524b82618 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.538142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:814b672bd49b7bcd70c827195a496134ff9e5a872cbe39ae13c063d800c1951a

Observation 11998235-261a-4a26-84fc-ee303e461380 · inbound

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model cites this paper.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.757511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.757511Z digest=sha256:89a6bbebc4fb468491fcd59e10c7ca98fc6c98278c5658e141aab8e868c17981

Observation 13d7ca2a-2279-434b-8f66-9cf72fd94b38 · inbound

CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech cites this paper.

CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:37.652047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:37.652047Z digest=sha256:fa417d337a98d4ffe9ebeea072506fe2644fd450ed932c3aee1d07fc54d11980

Observation f1241b30-69e3-4f52-90d6-97ee7d2b37d4 · inbound

Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance cites this paper.

Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T21:54:08.485720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:54:08.485720Z digest=sha256:f7281c64ccbe541e026af426c823948347dc0504973c4ea48844fcae330ad66d

Observation 78a7e2c6-acca-45c8-adee-6b6c05fad01a · inbound

OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching cites this paper.

OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:30:50.848560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:30:50.848560Z digest=sha256:523f4925b2229e21ce7ad3631dbd0c01c537ae2b7fe4a3013cb76914ab80bc77

Observation 2824899f-324d-4976-ad1f-b2f429daa910 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.474301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:a391b3d7ae5e048a543d43ca0b8b3cff6393968f5e9c9aebd16ce943384f2218

Observation b835ee00-fb8d-4e72-8092-748c120d4af8 · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.310389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.310389Z digest=sha256:587fab2b38600ff4f290e0a4e4f831f8f2316a17751faab477ac1e92d5a7a376

Observation 88b84e91-c007-4a5d-9186-1ade72b1f353 · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:29.295374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:29.295374Z digest=sha256:9057bbbd5646f86c2077c4f7945f356c7c500ae6f7ae707cb3a2e05bcd2a3e21

Observation d9ad5e0b-6bdc-447e-a44f-faaf8d83c312 · inbound

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching cites this paper.

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:21.098815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:51:21.098815Z digest=sha256:29aa5e6896e3d62eb9634c7a91ff6b07cc37089ec2939a824aa0ced063006539