Pith. sign in

Paper Citation Record · LEDGER

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

As of 16 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 4 inbound Pith citation observations for arXiv:2501.04644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04644 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:32:17.120091Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:11.587593Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:49:24.559099Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2b1e617-6d4d-4a45-aa9f-191bda070c2d · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.062098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.062098Z digest=sha256:9f6294bc33be9e1a9ba0f31a4ff1d4cd1a2c89049f4b4f2f5735691e2ed59759

Observation 49384579-9107-4709-b7a5-bcd2bcc0c23b · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.086034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.086034Z digest=sha256:c3bd2648b9f7839d938c9d32dbdec1da8a3fcb0bb79b0a2369e0cdb7b1165ed5

Observation b20f6d86-58f1-4b03-b427-5fe86fac6322 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.090803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.090803Z digest=sha256:823aff1979b54289afa6fa9d919a6009e7e9c712fa68db2958821af453aa5e8d

Observation 87d9f7e2-56b7-46f9-80fe-ff1f47f62436 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.100232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.100232Z digest=sha256:a1ec6d32c970ae11284efb1cda33e0b287fc3767c1f67afe137b680551142423

Observation 579e16a7-2a35-4fc1-845f-3364f0e92b3b · outbound

This paper cites Kazuki Yamauchi, Yusuke Ijima, and Yuki Saito.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Kazuki Yamauchi, Yusuke Ijima, and Yuki Saito

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:17.351318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T21:32:17.105230Z digest=sha256:8d9fa9bf6871130b363e68394a5b09adc3cc52595eb5991f4b26bed91ccd6e98

Observation 555b47d4-9723-4506-aee1-1e1c24e5b16e · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.110274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.110274Z digest=sha256:e15a2c7732ef2f6b1f6e8505ca1468b9af046664c1fdc453ca3da6db255d74f1

Observation 5c433296-5a26-40d5-a1b9-1a81bfce6c5b · outbound

This paper cites In ACM Multimedia, pages 7513–7522.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts In ACM Multimedia, pages 7513–7522

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:17.334493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T21:32:17.115337Z digest=sha256:30de682e085e22401000298d7e1d965548d211e043e340c431f96d51694dd9c8

Observation 6c7ba811-c705-41a7-9356-61b26b9b4c27 · outbound

This paper cites fast speaking rate.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts fast speaking rate

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:17.318179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T21:32:17.120091Z digest=sha256:e7493497dcc183a42918a2a5627373ecc716622a7a5e834034fac6799d6fef15

Observation c54bc1df-086a-417b-84d4-0d471703f3ce · outbound

This paper cites Emotional End-to-End Neural Speech Synthesizer.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Emotional End-to-End Neural Speech Synthesizer

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.071350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.071350Z digest=sha256:4a9823c344ad3ce93a1d55f7b50fbf97d9b607c66d0cdea9233218a9d44cafc8

Observation f8a02476-6325-48d8-8033-4140a9d3bf73 · outbound

This paper cites an unresolved cited work.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:17.367540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T21:32:17.095521Z digest=sha256:bef0555dcfa8b3613b4ef033b986dc189655e98f45d03d35ef2c21c741f76094

Observation 480d6631-a213-4098-a07e-240c853555d6 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.075729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.075729Z digest=sha256:93ee395d25986f0cffb7903daf7bcac20b82a070bf74171d3f0898a40c51e76e

Observation d2d2c83e-cb37-490d-9f91-737d913f5b85 · outbound

This paper cites In ICASSP 2023-2023 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts In ICASSP 2023-2023 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:17.383143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T21:32:17.067242Z digest=sha256:0575d536d5a6d645cbbc40dd9577fe840d82e7fb4c787cb383714934167750ac

Observation 3e5a84b6-e6bc-49b4-8250-4be61db26bdc · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.057150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.057150Z digest=sha256:c682550f9065d7e45874caeb2a9db7245c382a7db53b44069cbdbd0dc0a8802b

Observation 0632dcb1-5a7c-42aa-82bc-85664d44cfd6 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.050955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.050955Z digest=sha256:99aed2ace61383a99746c6d5f9f1b4ffc11c1a0f638d06b32b3a05a030cd7804

Observation 215e5912-a577-4fab-9ca8-de53817efa26 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 7771

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.080584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.080584Z digest=sha256:bec19f79affeb95fcc81c5adab7af84cbdce1d74c86d0e1cb6609e38e50db349

Pith citing papers

Observation fa975243-39c4-4a5b-8925-b279af81baf4 · inbound

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers cites this paper.

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:11.587593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:11.587593Z digest=sha256:686953c18389ced19c6785be10c33c88e8d3ec6a17cbcde920b686c588b96898

Observation ea35b21e-fc12-4662-bab0-88f46af95e88 · inbound

JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles cites this paper.

JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:54:45.973944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:54:45.973944Z digest=sha256:d5c52948a80e56abbf85f25946384f35b26b014e5e425e90a81442aa787fd05e

Observation ec264d95-654f-440a-9c15-2b2f25920f8a · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.065145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.065145Z digest=sha256:d2341b9a42297d7d19da3f494e5b4ac5697277cf0ff3e9e8f4569ceb6ea8de31

Observation 6a402457-476e-405b-bc98-8b35a98c947a · inbound

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech cites this paper.

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.560693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T19:09:33.605232Z digest=sha256:8fefc15dcc86a2a38cde41170ed07e3ceecf3198f240cdc66f5dece519ca1fff