Pith. sign in

Paper Citation Record · LEDGER

AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2010.11567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2010.11567 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:35:12.394750Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:49:38.607429Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8e2bd706-61c6-48a5-89eb-acf0618e489f · inbound

Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection cites this paper.

Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T04:35:12.394750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:35:12.394750Z digest=sha256:163c359cb57c45ae2caf2c5f4ffc02fbe0a8c0587d40dbb6bc10cb1cd72a4fab

Observation 080ee40c-76b7-47f6-b127-d023a047e778 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.038704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:da0289fdb30d1df33215b2a99519b1f39fc441c54ea678a66edaec8b334b6a12

Observation 72b3b378-ef9c-4449-8708-5f8c65c46d46 · inbound

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval cites this paper.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.239809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.239809Z digest=sha256:06024bcd47e4a2018f10f785e49825e46bcbf7a2428b6726b12f948bd4acb0ea

Observation b1404a95-a522-4169-adbf-291624ac1cc5 · inbound

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training cites this paper.

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:56.536930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:56.536930Z digest=sha256:a7e0077bcada0a14e33b30d98b280878243f380b4fee9aaef3ef4db0fc3c1b58

Observation a99b368e-9179-465a-ada5-1305950229a7 · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:21.272793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:21.272793Z digest=sha256:cf83313c979bcdd99b55355508b4d763b42be77424c8dfdc2ee127ede2c54207

Observation ae9dbea3-a6f2-40af-abba-98735084747f · inbound

UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling cites this paper.

UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T22:07:46.302600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:07:46.302600Z digest=sha256:1528f233115cd2362421581951823e8a867fc5520933eb6996ee8f9392131e84

Observation 9145723c-61f8-4897-9a5c-95f581107e06 · inbound

SwiftF0: Fast and Accurate Monophonic Pitch Detection cites this paper.

SwiftF0: Fast and Accurate Monophonic Pitch Detection AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:31:05.595288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:31:05.595288Z digest=sha256:971d4b082c97802b172a5a727bbbb2cb5d93bcface134841bf8f0d539297943e

Observation 3fe8a6d4-6c3c-4f6e-a61e-b88ac4cb3a0d · inbound

DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners cites this paper.

DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:53.149135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:53.149135Z digest=sha256:446fde46aea3c4533d668e0f509286925d5ff3f9b32bdae8caf237d9d40efb58

Observation ede1d017-df78-4cc2-bed4-1c84314f5094 · inbound

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance cites this paper.

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.785216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:34:15.220263Z digest=sha256:982b4ef9ade8b1c8a7930b77f8babda066f7930a54f3f975bce7f9a409f3b3ef

Observation 59b7d247-258c-4e58-81c3-665acdc0a55f · inbound

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement cites this paper.

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:59.187097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:59.187097Z digest=sha256:a03bb399c5d5db8ceac15636bcf34c60922d6c74456c16d59d1d3f6242717837

Observation e9434d92-ebe4-4f3f-b63e-32a431447960 · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.807719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T20:34:50.539351Z digest=sha256:5bbf1bafddb59339e6e370c638d24bd403cc94a7888d7be9cc4dc3a9c58efc4e

Observation 660a81a9-150a-4a4f-85c3-0cc6b8fda1c1 · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:33:13.615847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:33:13.615847Z digest=sha256:3aa07978553a1c5aa77fcb9fed483ede860cdf43607dc8f6cf779af4a4f6c346

Observation 382d7016-5903-49f6-8b9b-bb98d1d4fa4f · inbound

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation cites this paper.

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:01:07.558745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T17:00:59.421441Z digest=sha256:cd535d204b6705dd7da883656aef02498d29d864ff2cfda55947c16c8c9f1308

Observation 99101ee1-bdd1-40b3-8aad-c3249b48d02b · inbound

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan cites this paper.

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.656870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:40:10.657590Z digest=sha256:51ae9d8f233a6980d5cdaa1f332e6cde2ea2eca0d87b76c68e47e2b74a9b9636

Observation d9f90c94-f0b9-4589-90ee-a3386cf88e91 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.000225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:c70310294ab2cd58cc3aec4743af27d212a97e21629927dfbb566a6c5ead4bc4

Observation 5c0903ec-2150-4288-8a05-51b1f7ae5c0f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 129

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.944674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:1ccd576fd09c4b84bde330679e4978f1fa64fc53703e604e6275dfd20cbb47e1

Observation 9e17326b-787e-4924-a8f8-91a541982b3e · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:00.259741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:78e03c3d49b9d7ebd0e0ceb87c8975a57c5d6f9c66d97c102269bc89f6802db5

Observation 1436a86b-d04e-4db9-a1a8-371b85b70769 · inbound

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing cites this paper.

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:06:23.940166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T12:54:20.815371Z digest=sha256:c518d577cf4e962bfcf8e86359e1b3ea1d3c24058b97df433a8408d111e4c374

Observation 69a15ad5-1818-4586-ad04-70fbff2f53cc · inbound

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding cites this paper.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:56:51.761438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:287e8fcbbe0409df9813fe5fbe427800c76baf052d79eaba762f23f277b1898b

Observation b89d8ed0-6b2e-4c88-b7bd-05fd353eac66 · inbound

ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding cites this paper.

ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:47:44.574399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T11:51:25.305324Z digest=sha256:88b31a3202cf0ca2ded5f40b589d8585d027e82052d6bdf7e697c224f507dcf9

Observation 2e069813-4b96-4c0a-957a-6beeb27708c7 · inbound

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective cites this paper.

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:12.259675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:43:35.279033Z digest=sha256:2f504faf8648167dc582fc14d8347ee2d420d8de095bff2a0cb6265029b420f0

Observation d58346c9-cf64-4b04-90ef-49a30faf9bf1 · inbound

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction cites this paper.

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.751682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:07:31.795522Z digest=sha256:5033476ad4d31678c7aeef08616ae0ab8942b5df89ec7ca470720a6688a2f184

Observation a949a7e9-5d13-464e-a3a1-a3f94fe8517f · inbound

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation cites this paper.

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:38.609190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:53:34.428544Z digest=sha256:c040d28ec3137f73009098ac16eaaac2850847e7075ad906409460f5e2290023

Observation 7f9b7f3b-fe35-4c58-8aef-be0e603c8e3f · inbound

Teffic-Audio: Tell Fact from Fiction cites this paper.

Teffic-Audio: Tell Fact from Fiction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T10:28:23.043310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:28:23.043310Z digest=sha256:5da5ad9ad9ea167a47113702818ed63359cf7b17504ba1647d1b00682b5d11e3

Observation 38a312d5-8ddd-48a6-901d-e15527f0d233 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:27.881861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:27.881861Z digest=sha256:f1239ae9f72b39819334d048332b6fd7a72d900c36f8696432e1eb28786dca55

Observation a95b1244-7755-434f-a7bb-b75e33b5c60f · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.871653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.871653Z digest=sha256:0ea38a65be119934a99d0c0bd73a032d3890af2c6f7c9181556b8d2853a197e7