Pith. sign in

Paper Citation Record · LEDGER

Natural language guidance of high-fidelity text-to-speech with synthetic annotations

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2402.01912.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01912 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:43:08.790672Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 381e5d4b-ee52-4d40-b293-c6eadfcfbea1 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.598879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:05f1e9bdc2b4f0042f2d632eb0be3ea4da67ed000d979d03a81f152420dc3677

Observation cd010cce-1125-4d87-ad1e-da6c7c2af2ce · inbound

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis cites this paper.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.790672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.790672Z digest=sha256:3a624cf4b149ee3df414a49c4176afdc0dff4ae0e58e45fd3e1c26a6d1e4ffaa

Observation 1222f66f-e60a-4da6-a341-18a4d700dfa7 · inbound

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits cites this paper.

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:51.142199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:51.142199Z digest=sha256:6b94c061fed81a39161ecd739b7f3664cf11581582f5654906c6ef3f4880156b

Observation db7c9d04-f6fa-4db1-9bc1-def38a7a896a · inbound

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations cites this paper.

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:42.812459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:42.812459Z digest=sha256:05c023a9967bfa1d1bc9c2367ec871704a227c6acd01208be09609cff26410df

Observation 1d643197-ab5f-49a1-96c9-e56bb18d5d6e · inbound

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis cites this paper.

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:04.837891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:04.837891Z digest=sha256:6523f1fc4b70aec698938d6dd97972d8ab0dd15f8465c5d7a72112d03be6c7ca

Observation a0230389-f536-413d-b453-4a4d8395e27d · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.191205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.191205Z digest=sha256:c5a1fea12f4650f0d2fc0f7cba08d6f0b0beebeebbb60db971f0a04c1e9db974

Observation de9f7812-7aa8-4c1b-ba5a-0257ba1e96db · inbound

Optimizing Multilingual Text-To-Speech with Accents & Emotions cites this paper.

Optimizing Multilingual Text-To-Speech with Accents & Emotions Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:29.502416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:29.502416Z digest=sha256:501a8dc2d9a72034111d5b60c9109560a89f5f7fbc2e7b50b2af154c337f6790

Observation fb386c66-e20f-43ce-a4ca-d2c323e05703 · inbound

MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications cites this paper.

MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:13.177749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:13.177749Z digest=sha256:4256aecae1c791ec3c341a53c19f290313bdcdfe6bb4c12d078058cbdc43962f

Observation 5c2cf6a8-f131-4464-984c-ec5e78ae53f9 · inbound

Multi-interaction TTS toward professional recording reproduction cites this paper.

Multi-interaction TTS toward professional recording reproduction Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.456914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.456914Z digest=sha256:1f64547d89ca4ded83d2dab46476b39bdadca2ff83c18d7bf0484b9979c061ca

Observation 9d1545d8-6f18-4bb2-92e6-f9bf768217e7 · inbound

SecureSpeech: Prompt-based Speaker and Content Protection cites this paper.

SecureSpeech: Prompt-based Speaker and Content Protection Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:36:07.484115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:36:07.484115Z digest=sha256:e5c64f7e19eb814ef81976646eb7ca61081c56f6d5718af7e0bf4e6841c47e08

Observation cdb68026-84dd-4f38-be9c-6b9913761e75 · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:33.855401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:33.855401Z digest=sha256:0d6aae03fda82e56126400f029b739101530c6a252acdc7e6bc64a0416f37886

Observation aea71d72-4db1-4369-9af7-bc31544e2993 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.214444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.214444Z digest=sha256:9ead9193b8e1738390262127177274316b218908cf9b553b7249601472bf1c28

Observation 1ce65cf1-7f38-4863-bf96-b730fdb28a2c · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:19.629984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:19.629984Z digest=sha256:19ac3ba73185d4c82d4c6a52c96ed8efe5b5d73200672a1b30f7320a9a3946c6

Observation 604d43fb-5ff3-4590-b857-a3b1e3df899f · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.453562Z digest=sha256:dc46a51a40da316d344001ba0bdab1e9b28d510f17e15cdc085a9ac00763c65e

Observation aeae0d3f-9346-482b-8379-4d91f7bfc692 · inbound

Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates cites this paper.

Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:04.179780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:04.179780Z digest=sha256:af0cdac8ef7671a2e01df656d4462e2e94d8df21338e55a0029664938e70188e

Observation 160ea297-49c7-498b-8f39-1db404dc9e29 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.136038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:ff01b1980d65b011c3d5240a6666383122fc59dbb55d09c846309d64419169dc

Observation 949c6284-1e5f-489c-ab16-b3c6e92f45be · inbound

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus cites this paper.

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:30:09.664146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T16:29:10.899659Z digest=sha256:205d3088b9a52dfdae183a6dd193bfd4098879dd0ac84667644f2f0547c87902

Observation 91288b67-7835-4021-acc9-f75b037dc95e · inbound

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus cites this paper.

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T19:24:56.523051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:24:56.523051Z digest=sha256:adb572a0e0400148cb9681bd5ac61c730334dc39bd699d3b783980055ae80010

Observation 5ff90eac-0152-49b9-a645-8644c62c3689 · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:04:47.244302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:0ec16fba8fc2f699cdcbe5f27a3a80672ed47761d04c00e260c4dafc08d45a19

Observation 4b13ef46-b70c-4ac7-87e6-6a8f1720f315 · inbound

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation cites this paper.

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:04:50.107872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T01:01:06.094276Z digest=sha256:f54ff338e1e47a47484a8341cce0a92a5b33be11caf84d9017bae3469ab8cbbd

Observation 8758a34a-8d0e-4f9b-b29b-0d694d21bec0 · inbound

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment cites this paper.

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:10:09.324177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T01:07:50.243903Z digest=sha256:f60096efa6b1f98161c780e6f855b60f3280a86ff55e5f850a6a8665eab0c480

Observation 12c7d0ed-e83f-4cc2-907b-800c1af311b4 · inbound

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs cites this paper.

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:18.771185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:20:35.333172Z digest=sha256:7ef33b0a8b443fd8e6a4ece38492ee63947f1f9fb863b1b29d81eb89ccb785b6

Observation b0b47817-d91b-4f24-b4e5-bb1081c9a2ee · inbound

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models cites this paper.

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T00:18:36.390365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:18:36.390365Z digest=sha256:431a21571112140884335cf9ff8e37d3b924ab044c8e262ef9104056321d272c

Observation d9f6ae9c-89d1-4459-9a1f-92ba7f33e61a · inbound

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech cites this paper.

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:27:04.903956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T23:53:55.385445Z digest=sha256:a1e8bb162b12ddca021ecec73f90a1008301014bfe112cecb239bd9c95a9a8dd

Observation f00501ae-1501-41ad-b754-7230a5da2bd1 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.727005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:ebe016a3a5a551e8be3173f4902ffe92ffad631af17d2de9345259dcda25bffa

Observation b681c0fc-bc11-41d0-9b8e-f75c196e9e85 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.790821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:c09c9532ac13fdb6b5b30c1bf643a4a7816614acac62d70a6059189091104c1b

Observation 3ad237b5-80b3-46d7-9f7a-40afd9551ec8 · inbound

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech cites this paper.

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.577822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T19:09:33.605232Z digest=sha256:c37eb9f221c1054c820146aacd2e67ad323ab0cafbd59cbc4ca582c1b795334f

Observation 4655fea6-e4ec-4da8-8bdb-f9354cff3497 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.552835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:adbe5675cc2bcef70fb4044c953808cc421e938098a701b5f1423460d35de3a6

Observation 7e2215f8-42e7-4a5b-9347-715b6a74b32e · inbound

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer cites this paper.

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:46:13.750265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:46:13.750265Z digest=sha256:eac828df3e8ce95ff876bd639236bdd1f94378fdfe391699887d62221c2beb6a

Observation 69f255a2-f5d5-4f26-a467-9c911d7b2c6a · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:26.024843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:26.024843Z digest=sha256:1b65d1383921b5aedab17b041dedbab0be82a1c09d584ee59589736a927873b1

Observation d33c18b7-0caa-4bd3-9a9e-2e5011ebca0f · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.524756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.524756Z digest=sha256:74bd961a7c87f1a25e85e3d54b7f677695cb2155fcdc650e4cdf98a089da8c6e

Observation 8851c956-d783-434b-9400-67f0ca5946d0 · inbound

Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study cites this paper.

Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:58:58.620852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:58:58.620852Z digest=sha256:d83a81ee6f05d3906b06179fd165a57053cf39f4633656266e41c362150f8871