Pith. sign in

Paper Citation Record · LEDGER

Multi-interaction TTS toward professional recording reproduction

As of 14 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2507.00808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00808 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:11:40.541402Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:11:26.568638Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:11:40.627696Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5356663-a23d-46d0-be41-e81f4b445693 · outbound

This paper cites Multi-interaction TTS toward professional recording reproduction.

Multi-interaction TTS toward professional recording reproduction Multi-interaction TTS toward professional recording reproduction

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:11:40.632803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.568638Z digest=sha256:014713c00134eb9bd4e4b8db0bdd6697fda3747d3887dd22135815eee14f2772

Observation c49bf936-37ec-4847-a415-fd790f86d472 · outbound

This paper cites During the recording of this dataset, a di- rector iteratively gave acting directions, and the voice actor then reflected the given directions.

Multi-interaction TTS toward professional recording reproduction During the recording of this dataset, a di- rector iteratively gave acting directions, and the voice actor then reflected the given directions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.243995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.576860Z digest=sha256:3a81b64beb68ded8eed68d41ee23c7b5741f666d0ce3c59e0a0a8efa3ff9331d

Observation e1c22324-83e5-42e7-bef6-bbd725d26ea2 · outbound

This paper cites Speak more brightly,.

Multi-interaction TTS toward professional recording reproduction Speak more brightly,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.232486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.591023Z digest=sha256:2a215905b438c08ebaa37a8dca51d02237b333e7e43866a004d1da530383fd55

Observation 7d97edb3-756a-490f-b9ba-e537f8a8edae · outbound

This paper cites In- sert a silent-pause after this word,.

Multi-interaction TTS toward professional recording reproduction In- sert a silent-pause after this word,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.220883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.605668Z digest=sha256:841f9d9d72eef535ae2c2cc1995fc7284b39134231d4717eb62e71f8b6eaf52f

Observation e59eaec2-69c1-42b3-8cb5-86b257961fb5 · outbound

This paper cites Follow my example.

Multi-interaction TTS toward professional recording reproduction Follow my example

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.209689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.623467Z digest=sha256:b30be8938f30dfbfc167b3dfd07ad6e3ff103c9a5df84f5eb7c7c9fcfe002b9e

Observation 0bc539be-d5a5-48d3-9873-1a8d7fd58242 · outbound

This paper cites The Guideline for TTS Speaking Style Classification.

Multi-interaction TTS toward professional recording reproduction The Guideline for TTS Speaking Style Classification

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.195100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.642547Z digest=sha256:6b355305278f1233c5c3dcb8b018a25b12f30c13c600e56b74c0f4dc7e940565

Observation 589bc8ec-64dd-45fb-b977-05653a3c6f3c · outbound

This paper cites Dataset We used three Japanese 22 kHz datasets: interactive, non- interactive, and large in-house datasets.

Multi-interaction TTS toward professional recording reproduction Dataset We used three Japanese 22 kHz datasets: interactive, non- interactive, and large in-house datasets

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.182706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.663023Z digest=sha256:adbe2f9c2b900d13ab58145e9af46cc2881bbe346a4adde159b2a90944792235

Observation 2255d797-b963-460e-befe-03f7b0d40e36 · outbound

This paper cites an unresolved cited work.

Multi-interaction TTS toward professional recording reproduction Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:11:41.170322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.683475Z digest=sha256:eb439946cf7fdcb1cb9eaa4af5428007767c143927facb0f8a917c95808170f6

Observation 53851e5c-e299-44a8-bb96-0a63fe1a2e24 · outbound

This paper cites an unresolved cited work.

Multi-interaction TTS toward professional recording reproduction Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:11:41.159357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.694379Z digest=sha256:02bf2e3b645333b3e9ee23d19468aaf70251689924acf9216b859028ad278179

Observation 98649d8d-f586-442b-a582-b690c2f38fc5 · outbound

This paper cites The loss function and learning rate were the same as in the previous step.

Multi-interaction TTS toward professional recording reproduction The loss function and learning rate were the same as in the previous step

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.148255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.703922Z digest=sha256:96221712abec58bce355838897ee912eb33f7f27e7c4041b272e1856a3e1536d

Observation 84dc560a-abc9-4044-bb90-bc8db27c61d3 · outbound

This paper cites The training step was 10K steps with AdamW optimizer [41] with 4K warm-up steps.

Multi-interaction TTS toward professional recording reproduction The training step was 10K steps with AdamW optimizer [41] with 4K warm-up steps

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.137813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.715018Z digest=sha256:3364e3249559946aff0f2dc374ed3ef8be15a29dfee2dc9c218c86879e7cb0f6

Observation 62883a09-b4af-4998-81d9-34b4d500cd6f · outbound

This paper cites Overall alignment only.

Multi-interaction TTS toward professional recording reproduction Overall alignment only

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.125676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.725122Z digest=sha256:f5ba4e821d0760db500aa78b448587f520432a7c03578f39fd98d0e36a84fe8a

Observation 30701ae4-0d7b-4e3d-bc99-85fcfd78448c · outbound

This paper cites at the beginning.

Multi-interaction TTS toward professional recording reproduction at the beginning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.115057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.736074Z digest=sha256:7eaa0c01fe42e27cc85ebc93b1d07c5e2ddd41be14a2fb3b170dc712c537b5ce

Observation 11d0c593-3556-437e-8097-2fccaec969bf · outbound

This paper cites Subjective evaluations demonstrated that our proposed method achieved iterative style refinement that matched the user’s di- rections to some extent.

Multi-interaction TTS toward professional recording reproduction Subjective evaluations demonstrated that our proposed method achieved iterative style refinement that matched the user’s di- rections to some extent

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.104855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.749700Z digest=sha256:934616ea0a6f3443fc10c9d2f2aaaf1692f908bb61782fde729467cc89e26fb8

Observation 163bc918-16bc-40fe-a3db-a4eac5a8a983 · outbound

This paper cites The relevance of trial-and-error: Can trial-and- error be a sufficient learning method in technical problem-solving- contexts?.

Multi-interaction TTS toward professional recording reproduction The relevance of trial-and-error: Can trial-and- error be a sufficient learning method in technical problem-solving- contexts?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.092369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.754699Z digest=sha256:a1ba72871663c56a3b75dd985e7af0e9f7729ec9d03e97743aa5472affd268fb

Observation d8bf1404-3233-4bf1-9700-f8452d280a5c · outbound

This paper cites From da Vinci’s flying machines to a theory of the creative process,.

Multi-interaction TTS toward professional recording reproduction From da Vinci’s flying machines to a theory of the creative process,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.079650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.762988Z digest=sha256:cc8ff67b1c6b6e144637d5b17b05ab2257d3fb9e38597e2525fd5e421a79b220

Observation 1f4e05c7-a4fa-4026-a76d-176b02ac7acf · outbound

This paper cites What are the stages of the creative process? What visual art students are saying.

Multi-interaction TTS toward professional recording reproduction What are the stages of the creative process? What visual art students are saying

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.067214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.806058Z digest=sha256:c1616216ca1c68f22a313c37f3e4ba6ec8db99c6157fd63e8ca4c18ecedb82b9

Observation af7aab1a-6d09-4f31-87d0-ce2ac2493414 · outbound

This paper cites From page to stage: The director’s interpretation and picturization of a script,.

Multi-interaction TTS toward professional recording reproduction From page to stage: The director’s interpretation and picturization of a script,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.056318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.863549Z digest=sha256:6fec0342892222171267187baa7a74dede66069ceea271b51a1462cca3d352f0

Observation e678791d-f81a-4aac-9e6b-68544a57c169 · outbound

This paper cites Showing and telling—How directors combine embodied demonstrations and verbal descrip- tions to instruct in theater rehearsals,.

Multi-interaction TTS toward professional recording reproduction Showing and telling—How directors combine embodied demonstrations and verbal descrip- tions to instruct in theater rehearsals,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.043331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.952999Z digest=sha256:28dcd4a7327de72f9781e2327941715eba7be5c189c95024138f92939c44635b

Observation 6cf3832f-cff7-4458-a084-a642fbc34abb · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Multi-interaction TTS toward professional recording reproduction Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:27.044864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:27.044864Z digest=sha256:9f5a0ce5c1dc0a0edc1837046ef878ba18dbc3e5cf07c61927a351a64bd98275

Observation e58288b0-a4b1-4f92-a2ef-a146ededeeb0 · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Multi-interaction TTS toward professional recording reproduction High-resolution image synthesis with latent diffusion models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.030672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:27.144346Z digest=sha256:093b007e55a236ca60f34f6b283c7393c34e1aa3ed0f52820eca731f230ac486

Observation aae4d33f-0de3-41cc-b023-a8f231788bf9 · outbound

This paper cites Photorealistic text-to- image diffusion models with deep language understanding,.

Multi-interaction TTS toward professional recording reproduction Photorealistic text-to- image diffusion models with deep language understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.019811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:27.271811Z digest=sha256:34a158fc865a27a9e3e7b692ec080375f343c6b8ca14170fae3a873df21342a5

Observation 0a6f5387-cb32-4563-8ea8-44b89c4a90a1 · outbound

This paper cites Program Synthesis with Large Language Models.

Multi-interaction TTS toward professional recording reproduction Program Synthesis with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:27.358467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:27.358467Z digest=sha256:bee101a5915da645d493017f977d51a42accc0c1cec950a0628b8d5f36d0e20f

Observation 5ddfe0d6-1d71-498e-b9b0-ef9c9ae37a37 · outbound

This paper cites CodeGen: An open large language model for code with multi-turn program synthesis,.

Multi-interaction TTS toward professional recording reproduction CodeGen: An open large language model for code with multi-turn program synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.006962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:27.459503Z digest=sha256:7b66589fd9b566cb86df6e864a23b2ecede4160631afd76d5dc0ed7a3e8c7ef5

Observation d212c621-fc9d-4f65-a863-63aa3f829877 · outbound

This paper cites Training language models to follow instruc- tions with human feedback,.

Multi-interaction TTS toward professional recording reproduction Training language models to follow instruc- tions with human feedback,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.993049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.323645Z digest=sha256:dc1a30836a8dbc5da90188ea1157358285d24d5de2d1535cae0cffbed272aad8

Observation 75703a4b-a7e8-464c-a830-477c1c0ca0e4 · outbound

This paper cites PaLM: Scaling language modeling with path- ways,.

Multi-interaction TTS toward professional recording reproduction PaLM: Scaling language modeling with path- ways,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.980148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.341138Z digest=sha256:410514df4552b5e14524fe583adacecd551c3c921544b6f20f14ef713f1a9f57

Observation 8b928b82-e0bb-41eb-b8de-269c3720e067 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Multi-interaction TTS toward professional recording reproduction A Survey on Neural Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.364754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.364754Z digest=sha256:e5364d7e25f9654488e8b35ea1a273503e8a9814877753a46c03c453840e1ab1

Observation 229a5f78-d42d-43a9-9874-757d67685f5d · outbound

This paper cites A review of deep learning techniques for speech processing,.

Multi-interaction TTS toward professional recording reproduction A review of deep learning techniques for speech processing,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.391082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.391082Z digest=sha256:95c2d53d6ad82101537edca3a8d1b32767d7e7d76c01b39c81fcad10b0abe74b

Observation 4b376339-a22d-4f5a-b305-0dcea196f62a · outbound

This paper cites Model archi- tectures to extrapolate emotional expressions in DNN-based text- to-speech,.

Multi-interaction TTS toward professional recording reproduction Model archi- tectures to extrapolate emotional expressions in DNN-based text- to-speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.961913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.408756Z digest=sha256:09225874ca111e77a4b4759052feecf31b04f5df93d6b2be4c1ac5ed33479501

Observation 4da98607-1c1e-42f2-a736-10a0cfe4ae89 · outbound

This paper cites V oice puppetry: Exploring dramatic performance to develop speech synthesis,.

Multi-interaction TTS toward professional recording reproduction V oice puppetry: Exploring dramatic performance to develop speech synthesis,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.948640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.417111Z digest=sha256:5288d27d481f0a4f5d3da0d9dfc3e877692189ba286eb706665734a28dfc756b

Observation d7d8a15b-bdbf-4a10-ad7b-6e9bb3c7339a · outbound

This paper cites V oice puppetry with FastPitch,.

Multi-interaction TTS toward professional recording reproduction V oice puppetry with FastPitch,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.936595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.420350Z digest=sha256:bd23c1952152bbe9affb9d7bf71c47b194072f978490733173b4589f5327bc22

Observation 6a885f33-c101-4bd3-8e7c-a2a3de59c71d · outbound

This paper cites Style Tokens: Unsu- pervised style modeling, control and transfer in end-to-end speech synthesis,.

Multi-interaction TTS toward professional recording reproduction Style Tokens: Unsu- pervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.925522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.423738Z digest=sha256:f41734eb15aa0343d29342a40b9c9f40bd9d3bd4f94520b3f761f5bdce1cfccf

Observation 0858633e-7ff2-4f3a-909d-9a0cb74122cd · outbound

This paper cites Robust and fine-grained prosody control of end-to-end speech synthesis,.

Multi-interaction TTS toward professional recording reproduction Robust and fine-grained prosody control of end-to-end speech synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.914140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.437462Z digest=sha256:381f2998f48f43b03a36e448736e74e33b726a9f2cecc76be31fcee166b77e6c

Observation 82dd43d4-0a36-4e58-8ede-1133e2d234d4 · outbound

This paper cites Fine- grained robust prosody transfer for single-speaker neural text-to- speech,.

Multi-interaction TTS toward professional recording reproduction Fine- grained robust prosody transfer for single-speaker neural text-to- speech,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.903959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.442586Z digest=sha256:15bc661385d84264e500e9c2a7cd7e0b0d72139a8eb373c959aae9db46f682c3

Observation 92a808f0-8fd2-4add-a03d-b91a9555accb · outbound

This paper cites Daft- Exprt: Cross-speaker prosody transfer on any text for expressive speech synthesis,.

Multi-interaction TTS toward professional recording reproduction Daft- Exprt: Cross-speaker prosody transfer on any text for expressive speech synthesis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.892893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.448323Z digest=sha256:377b8f5264a7f1bd12b7e58ae332223ce4ecab363d088b7ae71fd685a30f17cd

Observation 3e7cf286-77db-4070-8541-5dba1a3fde25 · outbound

This paper cites PromptTTS: Controllable text-to-speech with text descriptions,.

Multi-interaction TTS toward professional recording reproduction PromptTTS: Controllable text-to-speech with text descriptions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.882082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.452192Z digest=sha256:32213dfdfb2b3f12fbb8b5e79ee1da04a907c321b3fda1a7f619d093942b497a

Observation 5c2cf6a8-f131-4464-984c-ec5e78ae53f9 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

Multi-interaction TTS toward professional recording reproduction Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.456914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.456914Z digest=sha256:6ff5c9cf9e33229bc40c8f8a2d9751796321b7070386c90cd6e71100594336e9

Observation 8e89fee7-5269-4fea-9a9c-9fcabc87d058 · outbound

This paper cites VoiceCraft: Zero-shot speech editing and text-to-speech in the wild,.

Multi-interaction TTS toward professional recording reproduction VoiceCraft: Zero-shot speech editing and text-to-speech in the wild,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.870810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.460077Z digest=sha256:698bd56e0920dadf4d6a78050ef3b61a8a5ac1bcb616e57506790400a11c9dac

Observation ee147f85-f3dc-4ce7-ab12-bd22785346eb · outbound

This paper cites V oice at- tribute editing with text prompt,.

Multi-interaction TTS toward professional recording reproduction V oice at- tribute editing with text prompt,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.859744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.463626Z digest=sha256:5199ea75a54a3b3617750d5be1e3fc9df5c12521b30dd9262e6a685c5c04459f

Observation c3f6fd9f-63d3-4cb1-8b9f-21642b264989 · outbound

This paper cites The guidelines for TTS speaking style classifi- cation (IT-4012),.

Multi-interaction TTS toward professional recording reproduction The guidelines for TTS speaking style classifi- cation (IT-4012),

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.848015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.466780Z digest=sha256:bf97687c0f1443666063e69811fd7b0593f28c8907d9a68486cfec022aa10ddf

Observation 1c9e30ab-a281-45c9-baed-3c15a4a6478a · outbound

This paper cites Zero-shot text-to-speech synthesis conditioned using self- supervised speech representation model,.

Multi-interaction TTS toward professional recording reproduction Zero-shot text-to-speech synthesis conditioned using self- supervised speech representation model,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.836114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.470345Z digest=sha256:130ac9ce3a3c948b05993c373ad32288b52f597f4fe9b434e80e21017a83f81a

Observation edd0a729-303e-49e2-817c-e4fd0189a9cc · outbound

This paper cites Why does self-supervised learning for speech recognition benefit speaker recognition?.

Multi-interaction TTS toward professional recording reproduction Why does self-supervised learning for speech recognition benefit speaker recognition?

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.824285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.474432Z digest=sha256:af5b2ff626e90b21423c4c407a8f96a03598b2acf37825ea055c03d44a59f1c4

Observation 90c31184-97b6-4947-8cf3-5ad8aea0dadf · outbound

This paper cites Feed-forward networks with atten- tion can solve some long-term memory problems,.

Multi-interaction TTS toward professional recording reproduction Feed-forward networks with atten- tion can solve some long-term memory problems,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.812387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.478099Z digest=sha256:f21d6f595ecfc5e7c558935d4c09c73811ff5bfa2915052d5e262ea6f082b06a

Observation 3f58c8c9-6a46-4d3d-8bef-7e44cc592c25 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer,.

Multi-interaction TTS toward professional recording reproduction FiLM: Visual reasoning with a general conditioning layer,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.799080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.481854Z digest=sha256:2c4f18bf1d6f362b51ae74c2d7d9996f562fbc3c63f64a206186afd680b24bb8

Observation a142c6f3-6c21-4885-9782-5769efddf13f · outbound

This paper cites Learning alignment for multimodal emotion recognition from speech,.

Multi-interaction TTS toward professional recording reproduction Learning alignment for multimodal emotion recognition from speech,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.786499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.485594Z digest=sha256:184a110c09457338df8fae9b722ab5a79b151df98736b63d6fb619fac3a02818

Observation 6f8acf10-56df-4593-9a55-93ff0820cb56 · outbound

This paper cites Multimodal cross- and self-attention network for speech emotion recognition,.

Multi-interaction TTS toward professional recording reproduction Multimodal cross- and self-attention network for speech emotion recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.774085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.489331Z digest=sha256:9c9ca4c537f2b38c5fa5ba7adddf5c7cdaf68199690a8d2c3265160cd2308dc6

Observation 30949dc0-96c0-4ef4-b9ea-9bc9b9e9bf48 · outbound

This paper cites Hello GPT-4o,.

Multi-interaction TTS toward professional recording reproduction Hello GPT-4o,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.761622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.492151Z digest=sha256:8e7ddbd7db5d01935a73cc5c1820db9ee05437a6fce6a20d63379fbd4611380c

Observation 2356dee1-9d92-443f-97df-de095f883371 · outbound

This paper cites Rephrasing the web: A recipe for compute and data-efficient lan- guage modeling,.

Multi-interaction TTS toward professional recording reproduction Rephrasing the web: A recipe for compute and data-efficient lan- guage modeling,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.750736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.498128Z digest=sha256:277fbaa93bd80c89268e955da2a8afb950685b0b65a47da643df6ca1d5128ddf

Observation 04d3e439-1ad4-410d-a1c7-fdd0af902c2f · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech,.

Multi-interaction TTS toward professional recording reproduction FastSpeech 2: Fast and high-quality end-to-end text to speech,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.740424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.502456Z digest=sha256:76a25f19eb010edc3daaf3ec48fc8b63ea5cfbe93dba757bfcf7204df0b5c538

Observation 24aeb924-e40f-458f-ac95-331827345ec7 · outbound

This paper cites In- vestigating on incorporating pretrained and learnable speaker rep- resentations for multi-speaker multi-style text-to-speech,.

Multi-interaction TTS toward professional recording reproduction In- vestigating on incorporating pretrained and learnable speaker rep- resentations for multi-speaker multi-style text-to-speech,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.729800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.505360Z digest=sha256:2cdca9d3b6bdfa75d0211d88cf4901984fcac547267ab6d7129095e596612fab

Observation 0bc8d84b-6de6-49fa-9a4b-4c6bfcf054b6 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Multi-interaction TTS toward professional recording reproduction HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.508676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.508676Z digest=sha256:c5f5812ed1184eff793575d489e9ef5ba3486ba4e430f1b8e9e7fd3c4302456c

Observation 3dc66171-3503-43ed-a7c8-4a85e6d20c9b · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Multi-interaction TTS toward professional recording reproduction The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.512618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.512618Z digest=sha256:ee661f85b46d95b82e39126e271492542a8f06647a0c7f603d8f8958ce5d5cbd

Observation 2e60fafe-82e7-4392-80da-a62b9d224e36 · outbound

This paper cites Multi-speaker modeling for DNN- based speech synthesis incorporating generative adversarial net- works,.

Multi-interaction TTS toward professional recording reproduction Multi-speaker modeling for DNN- based speech synthesis incorporating generative adversarial net- works,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.713292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.515996Z digest=sha256:b192b9ff7f225d77937a9c77cede6726468126941c44b24f2bcf5c3193a75a4f

Observation ac610f13-f5c7-4140-b63a-55e47da9bec4 · outbound

This paper cites Variational discriminator bottleneck: Improving imitation learn- ing, inverse RL, and GANs by constraining information flow,.

Multi-interaction TTS toward professional recording reproduction Variational discriminator bottleneck: Improving imitation learn- ing, inverse RL, and GANs by constraining information flow,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.703451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.519705Z digest=sha256:12af3b8182f40c079d961d220521ebb0380066ab388d1e4b3211b097a6ec3ba7

Observation ddfa18e2-47f7-4b0b-8e0b-9cbe15bdf6c9 · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Multi-interaction TTS toward professional recording reproduction Decoupled weight decay regulariza- tion,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.523505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.523505Z digest=sha256:7a6849f06d12efc8595fc51a880e58000cc59c895eaa2b5d8e932ea3518a6109

Observation 03f70cff-87dd-4cf0-8dfb-4b54da726799 · outbound

This paper cites Expressive text-to-speech synthesis using text chat dataset with speaking style information,.

Multi-interaction TTS toward professional recording reproduction Expressive text-to-speech synthesis using text chat dataset with speaking style information,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.685250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.526777Z digest=sha256:b824e158b47e6f0820a76a4ef44f2b3c182613b32a708295302f12fba87847e9

Observation 79df0849-aaf6-4be3-923c-1d70f86d183f · outbound

This paper cites NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Multi-interaction TTS toward professional recording reproduction NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.673622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.530589Z digest=sha256:7f9fa2964c501fd5d57a609fc6b557c26eba42d9e4f2216bf6aace8b001537e1

Observation 55c5343b-e23d-4e33-94bd-556f25af4981 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

Multi-interaction TTS toward professional recording reproduction Neural codec language models are zero-shot text to speech synthesizers,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.663664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.534096Z digest=sha256:ff63232c21ce3f165bcad1c35976ff81d9c6c3e4c621e4ac5eaa04277b018db9

Observation 448d23c8-497b-4cee-9f0c-fb3ba3d0a156 · outbound

This paper cites FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation,.

Multi-interaction TTS toward professional recording reproduction FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.652925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.537401Z digest=sha256:2e3a7f9aa52fba78d8e1e5f53ceb997cc68ef34fdfb00242faff1fcb8cc23226

Observation 27e012b0-5b8d-47a2-a991-ebdd4448ff7e · outbound

This paper cites Toward verifiable and repro- ducible human evaluation for text-to-image generation,.

Multi-interaction TTS toward professional recording reproduction Toward verifiable and repro- ducible human evaluation for text-to-image generation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.643029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:40.541402Z digest=sha256:3406e82d054bd6b80ad4fd0fa980c40cc52aede8403aeef7305d80074712e382

Pith citing papers

Observation f5356663-a23d-46d0-be41-e81f4b445693 · inbound

Multi-interaction TTS toward professional recording reproduction cites this paper.

Multi-interaction TTS toward professional recording reproduction Multi-interaction TTS toward professional recording reproduction

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:11:40.632803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:11:26.568638Z digest=sha256:014713c00134eb9bd4e4b8db0bdd6697fda3747d3887dd22135815eee14f2772