Pith. sign in

Paper Citation Record · LEDGER

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

As of 15 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 4 inbound Pith citation observations for arXiv:2501.04644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04644 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:32:17.120091Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:11.587593Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:49:24.559099Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2b1e617-6d4d-4a45-aa9f-191bda070c2d · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.062098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.062098Z digest=sha256:8cbf8e6b2ae7e891b1d8e1bc7175e7c8122155726afd6666e139550aef354e06

Observation 49384579-9107-4709-b7a5-bcd2bcc0c23b · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.086034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.086034Z digest=sha256:ababe1ad7e7ebb86f0fcf34fa94fe3084a67324204f9ee846eac6c6cdfb82e9c

Observation b20f6d86-58f1-4b03-b427-5fe86fac6322 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.090803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.090803Z digest=sha256:f21f2f38446543aac68ae181f79f31e249980b5460545daf7312d3e618744b4c

Observation 87d9f7e2-56b7-46f9-80fe-ff1f47f62436 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.100232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.100232Z digest=sha256:8a103671252ee32609846dc9066aedcceba1a67b6e37c00ad2e6d63d9a9ea88c

Observation 579e16a7-2a35-4fc1-845f-3364f0e92b3b · outbound

This paper cites Kazuki Yamauchi, Yusuke Ijima, and Yuki Saito.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Kazuki Yamauchi, Yusuke Ijima, and Yuki Saito

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:17.351318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:32:17.105230Z digest=sha256:244b6d49da5e5a9a6139cdcaf1788fb7c2f0e2c2a28f12078df200b0b3dd30f1

Observation 555b47d4-9723-4506-aee1-1e1c24e5b16e · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.110274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.110274Z digest=sha256:e15a2c7732ef2f6b1f6e8505ca1468b9af046664c1fdc453ca3da6db255d74f1

Observation 5c433296-5a26-40d5-a1b9-1a81bfce6c5b · outbound

This paper cites In ACM Multimedia, pages 7513–7522.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts In ACM Multimedia, pages 7513–7522

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:17.334493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:32:17.115337Z digest=sha256:a0f167c7bf316645767375322f2843559e01af0baf9f5cf683a8cd81a6b7bfaf

Observation 6c7ba811-c705-41a7-9356-61b26b9b4c27 · outbound

This paper cites fast speaking rate.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts fast speaking rate

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:17.318179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:32:17.120091Z digest=sha256:d28143c7d91e0e73455cbad37a35e67db2c29445d6247808aba69eb2f8b949cc

Observation c54bc1df-086a-417b-84d4-0d471703f3ce · outbound

This paper cites Emotional End-to-End Neural Speech Synthesizer.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Emotional End-to-End Neural Speech Synthesizer

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.071350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.071350Z digest=sha256:4a9823c344ad3ce93a1d55f7b50fbf97d9b607c66d0cdea9233218a9d44cafc8

Observation f8a02476-6325-48d8-8033-4140a9d3bf73 · outbound

This paper cites an unresolved cited work.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:17.367540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:32:17.095521Z digest=sha256:333fa7d79ec5cac8625dbd16461d998b418f4ca4e11e8764987a4ba258170634

Observation 480d6631-a213-4098-a07e-240c853555d6 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.075729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.075729Z digest=sha256:93ee395d25986f0cffb7903daf7bcac20b82a070bf74171d3f0898a40c51e76e

Observation d2d2c83e-cb37-490d-9f91-737d913f5b85 · outbound

This paper cites In ICASSP 2023-2023 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts In ICASSP 2023-2023 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:17.383143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:32:17.067242Z digest=sha256:dbda31c0458f9abe756d9877c723cb75ace52c5791398c291e7d25c4c85d9440

Observation 3e5a84b6-e6bc-49b4-8250-4be61db26bdc · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.057150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.057150Z digest=sha256:d7aab00268e6489ac649832049335b7e0b5681b9a5b992a32f1dc38ebb5b355c

Observation 0632dcb1-5a7c-42aa-82bc-85664d44cfd6 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.050955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.050955Z digest=sha256:be96cbf94326293ef661f6f5b80977f52cf1a940d1845b936411209a1df65985

Observation 215e5912-a577-4fab-9ca8-de53817efa26 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 7771

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.080584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.080584Z digest=sha256:a3817e6cda009c6e45d568ca34806c7d6fb9d746affe1f310f3f963ba71361c9

Pith citing papers

Observation fa975243-39c4-4a5b-8925-b279af81baf4 · inbound

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers cites this paper.

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:11.587593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:11.587593Z digest=sha256:686953c18389ced19c6785be10c33c88e8d3ec6a17cbcde920b686c588b96898

Observation ea35b21e-fc12-4662-bab0-88f46af95e88 · inbound

JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles cites this paper.

JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:54:45.973944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:54:45.973944Z digest=sha256:d5c52948a80e56abbf85f25946384f35b26b014e5e425e90a81442aa787fd05e

Observation ec264d95-654f-440a-9c15-2b2f25920f8a · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:56.065145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:56.065145Z digest=sha256:5a499cf4e66160b468593b20eb13f4a7a9f32e404ebb340a58fe94a16c7cb9f1

Observation 6a402457-476e-405b-bc98-8b35a98c947a · inbound

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech cites this paper.

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.560693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T19:09:33.605232Z digest=sha256:e854c1a6530fdf8c8093e8ebb4985581d9f99c90b0f1806c45847098534f9831