Pith. sign in

Paper Citation Record · LEDGER

Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2407.05361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05361 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:18.525240Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:37:35.221767Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0847645e-1357-4457-be39-54962b7a6d8c · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.461429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:16ce39b0e858b3a1334438713d94d500c1a0f5320ef912178bed869637d4e92d

Observation 12b976f6-b36a-401e-b381-059329f04c31 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.525240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.525240Z digest=sha256:2c9f7b1cb969e2282f47f88842ceda10b3d859fab65abaea165947c88eba57b6

Observation f3fe19ef-6c80-4bbd-9835-209e6f266201 · inbound

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing cites this paper.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.212017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.212017Z digest=sha256:70e57d5f520af47d310c54a2632ee4df70a627e6748354b690420b8dfb996f8f

Observation aa6c3185-dcb8-4aa5-93ae-5dd8727a85ab · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.604407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:55.604407Z digest=sha256:2e3b166d995202dc6a56006f547ba95367a5aa5703ef7cf407db112d2e502f46

Observation 619ad11f-4d92-493b-ac30-edf1ab09139c · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.862140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.862140Z digest=sha256:5cc44b4e5de9d0fd666dac980c666b8aab7d11b4515d5ad125a981d2f99ecb49

Observation 9984cf1e-8e2c-41fd-8878-4c2e019d5754 · inbound

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding cites this paper.

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:33:09.346129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:33:09.346129Z digest=sha256:64e407ea139c99c330a73a5a09ec9c1742bd166a1394ee1a3ceaaf11baef8f38

Observation 89261381-9e78-4f79-bcef-a759ffdab20e · inbound

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech cites this paper.

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:52.444640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:52.444640Z digest=sha256:937d832a8a8c3ecf358ad27cba2837e36ab21de6655596a16ddfc372258b00d8

Observation d4a39cdd-a493-4d96-90f3-f2082407f8dd · inbound

Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification cites this paper.

Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:36:16.250981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:36:16.250981Z digest=sha256:f0627dd55c32ff613c64557c71cf59ae0bf2cdcf4eec5e2365ae37817744b136

Observation 2dd37568-2bd1-4ee7-bb30-6599b6d618ad · inbound

REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers cites this paper.

REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:51.083471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:41:51.083471Z digest=sha256:4b165ee097deb64e9eb4bebe2c836bf246b91078edd7b2490cf4b037732e2fb9

Observation 8d757418-74ce-4c4a-bd83-d8a955f0829b · inbound

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation cites this paper.

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:41:00.064681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:41:00.064681Z digest=sha256:9a043b459ea69930039334571fe52c30a664350c5a83534d30df53328d383a22

Observation 8623222a-d6e0-4eec-9ba4-f132b042c462 · inbound

Over-the-Air Adversarial Attack Detection: from Datasets to Defenses cites this paper.

Over-the-Air Adversarial Attack Detection: from Datasets to Defenses Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:38.516808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:38.516808Z digest=sha256:a97043ed2da717a14eefc4a9c20503f1189798bb87616b55302e2ca040c67b2b

Observation 67778385-97e8-4920-b666-a32c0c30aee5 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.656117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.656117Z digest=sha256:38d63bcc0ca4e45305894c09c36ce10f5a2e4b7e3a9ba93d2bdc45a0d85a38b6

Observation 74e8b097-9b56-4718-84c0-761387fba4d7 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:42.435007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:42.435007Z digest=sha256:e1fdc450bc7d1ec5107f6a142f9647b58f135817af29ea3c6f1427603bf8c04d

Observation 608f90e4-bc54-4082-943e-6bf1fa17f6bf · inbound

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration cites this paper.

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.893750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:15:17.804936Z digest=sha256:736c9f1cb25d420783f0b0f051dfe9c86d73d07f47c71eeb4412a049031acc8a

Observation e34932ed-ecdb-42da-a5f3-21cdaa1d3d9f · inbound

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS cites this paper.

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:16:11.062977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T21:31:09.399885Z digest=sha256:e1a32d1cb6153d15a6f1eb9d270d938df5b83a91af59591ddc3d2f3be0eab187

Observation d62537a3-9ff4-49c5-9a60-64ed195b9888 · inbound

BareWave: Waveform-Native Flow-Matching Text-to-Speech cites this paper.

BareWave: Waveform-Native Flow-Matching Text-to-Speech Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.223279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T15:14:03.681833Z digest=sha256:1d81d642c9ede2bcc5d7b85548669e4ee6fc3c83d280fe967c7c4497c170beec

Observation 7de01159-c38e-444b-aa96-f909a7d722ed · inbound

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders cites this paper.

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:30.042911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:09:36.845217Z digest=sha256:9db35754fdb6779bea2a232bb5400e56a316f2831cf3201e9e89f9d108f10ed6

Observation 1524d8be-3023-4edb-aeea-5ff6d095c286 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:47.191378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:3700324f209c95ea2286b85d0a0af5e0c0629d59e95ec5819d3ba0066d61130d

Observation dcfe1bb0-7916-43a6-ad3b-06d4a79bf3d3 · inbound

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis cites this paper.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:14.162197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:14.162197Z digest=sha256:5ccc5b64cd39d57c1bdfbe8021cb5c1e78e05fb0ca88079e17b3335132565187

Observation 5cca810b-9b01-4a15-9bcd-fb0ddc527fdc · inbound

Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness cites this paper.

Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:25:28.128920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:25:28.128920Z digest=sha256:dd49f44c7cdfc9daa55aa3f8bb9631671a42a165193a85082c8a33800f09cfd3

Observation 552f8c7f-7728-4854-8caa-0392f27ae99d · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.291105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.291105Z digest=sha256:334c304e295f0e27074fcb98f6a3dbab22d1e8e7ae62c69c9cb0d32069602f37