Pith. sign in

Paper Citation Record · LEDGER

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 13 inbound Pith citation observations for arXiv:2506.16381.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16381 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:32:31.021219Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:43.938774Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1c4c5413-a55b-420f-8993-e950f3df6c01 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.935359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.935359Z digest=sha256:e62f58be93c9f54e0aba8777b57660fc65f7a2ea51334ce6ee101e534545b71b

Observation 91eef112-d517-4018-9592-b53ce9e3043a · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems LLM Evaluators Recognize and Favor Their Own Generations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.939918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.939918Z digest=sha256:11acb417b885d14c53017d81fc1a98741d6be80e6dfc8078ac7d876f5269a056

Observation d03331a2-b034-4e69-b52b-507ea25d091f · outbound

This paper cites when saying …, raise your voice.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems when saying …, raise your voice

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.254377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.971061Z digest=sha256:4aa823d28b9a359b75d29dfec19336f7fbb50096b19b761afc194d0fa4142b08

Observation 5f200a37-6991-4fba-bf76-442b572898eb · outbound

This paper cites like,” “imagine,.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems like,” “imagine,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.185406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:31.001543Z digest=sha256:09e2be6aa0e5fa1d0c63fd371f4d42dbc4336eccd728ca4cd8db4d95d7ce405a

Observation 0db05b81-2181-44f9-8ddd-52823d3163cf · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.952653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.952653Z digest=sha256:1808dbb66458b44aacf5b9850ef905e5f58bb5c7022dd2f01c16fd009001db81

Observation e0b56e1c-db55-4849-9e07-15ac14b7bf04 · outbound

This paper cites high female pitch.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems high female pitch

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.301900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.956652Z digest=sha256:bcac276d2179f124f689350c4fe97dfcabaafbc10b08707819f0650dece37f3d

Observation f06d8854-60c0-45f2-9703-7a8dc91a2b40 · outbound

This paper cites hoarse,” “furious.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems hoarse,” “furious

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.289749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.974641Z digest=sha256:b74e9671b52a47ce3c8cfea51cf146dc230ee55ac87ab053e0d772133bfc5d80

Observation 3e853945-9f3a-45b8-bd27-fc6d7bd45681 · outbound

This paper cites • Use a rich palette of synonyms and idioms; employ different grammatical structures (e.g.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems • Use a rich palette of synonyms and idioms; employ different grammatical structures (e.g

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.278001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.978278Z digest=sha256:0be4c846a41c0157cf29fcb89c63b765de8107b4e1ae1637f81d8757fb9a7e3a

Observation 08a20cb0-8be8-44d6-86a1-0022dd76d1c2 · outbound

This paper cites Imagine,.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Imagine,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.265910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.981877Z digest=sha256:46dba769cd444ff7a4ffb9f13b038ba481d9221e6c921fe3eba8bb4c10467380

Observation 4a11d729-592a-4dd6-ae7a-f2bbe45bfe67 · outbound

This paper cites when saying …, raise your voice.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems when saying …, raise your voice

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.242921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.985106Z digest=sha256:4016b436957019745a060fd7d65d1de96eeb3c4223c0b4d5e6828d7e023da560

Observation 301e4e57-180f-4a97-b8b9-4d01c89b3a7c · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.232396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.988091Z digest=sha256:70570585678dc3c1845ab6f869f3f69c83b68ceb22be9702cd89d63fdf8ba8f1

Observation 7a02cb83-833e-472c-80e7-bfd981841e77 · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.221501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.991344Z digest=sha256:04f1a190f9bdfa01291c841fb10d2f3b69b4d0d815cde90e3e24ee43c67beaf1

Observation efc7ce45-d0a9-41f6-8a0a-b410d502896c · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.209734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.994683Z digest=sha256:2cb4b1b0821977b074edd28d2e9e81fd5e19d3afd65deebcd7b41717393e21ac

Observation 7675d19b-e74c-4928-bfc7-f8ed52a14284 · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.196676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.997735Z digest=sha256:f308afe3525957f9f4927ce63033e591a8b0029abe23163f39d884f6a21b1d89

Observation a05a3d84-339f-4ea6-877c-c13a9104167e · outbound

This paper cites general,.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems general,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.171745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:31.005980Z digest=sha256:c4204da0eed10a0eeaaaf359f9bbb8a0140538735037251e6fc025c5fbed0198

Observation c9d1562b-3952-48f5-8daa-02443c1c74bc · outbound

This paper cites Character + a single minimal action or speaking manner.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Character + a single minimal action or speaking manner

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.159846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:31.009393Z digest=sha256:d6692e7a31004c7d172422c62c803f852d632621013197498f8f856f80fab660

Observation 211fd641-cfb0-489b-b32a-60e1d9dbb735 · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.146783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:31.012990Z digest=sha256:390b97a548b71845257be54c2640d6e56bdb4ac06d96f26cdb908768b2f1fdef

Observation 3bb0264f-b38d-42c3-a195-c929f1072b7c · outbound

This paper cites All three instructions must differ and must not start the same way.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems All three instructions must differ and must not start the same way

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.134944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:31.016944Z digest=sha256:173b9672db4b759d15882e29ff30909fccc5ece3f666145a59fab4a8c141987a

Observation 3ac62693-adb2-4219-b1ee-9ac68704f33b · outbound

This paper cites female high voice,.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems female high voice,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.123074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:31.021219Z digest=sha256:a3d7e584e7f5b6ed35b6037f18d7c24a7e70ddb2789922889cfa1bc819ea26df

Observation d5e5367b-6e78-4ca4-b051-14f110120e45 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.944300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.944300Z digest=sha256:b67b7a7409cca6fb5ac3c962421b7f571d31de373debe2b9631496eda15787b8

Observation 266c7ed0-129f-4467-bd76-66023d4bf56f · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.930017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.930017Z digest=sha256:cf71f210b5936f1793f92a6be9940d6445c6cdcef8609faf6efa726021ed4511

Observation dd1aede6-4060-4a96-bf98-5a7372a53934 · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.316876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:32:30.948787Z digest=sha256:61a41e788eed7c545bd760968db8c2a3252ea5ead629a67941aedffa306b62d2

Pith citing papers

Observation af1056ca-f61a-48c9-af15-a2cdf23ec4e0 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.122236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:fb1d0e35c55723c90eb96f9150a756dea4a75381614dadd0cf0dc8eca0ee975c

Observation d868c2ba-6f12-4b3a-8397-71b4977185e2 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.213627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:35a3ee644b587b8647cb0970d3e6eba21c3230faf27edc9d4b39085ded90746a

Observation 88bbdbc7-9e49-4a19-a837-94f66f121876 · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.979580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:94868d838ff59dee38b67db5e3ba0c28518544ef4443dd874a76ec375dc1c308

Observation e78d9862-8b72-4ad6-bed9-2b81c554d029 · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:04:47.258265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:8c14d416fd4cb9717cabef46538015c7f0afd00513b0f600403c3d50d79bde00

Observation 19987e2b-a1cb-44cc-b390-ade3ae5f29da · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:249956a7cfc713b4dddaaf8ebf7000ea5db6f3960c51a6bbbad3fa6a4a0000fd

Observation b976e3ac-df8c-4e34-89b3-0513ac085e4a · inbound

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge cites this paper.

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:00:05.346556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-25T22:07:10.531784Z digest=sha256:22e332bda8bf39ae2f1de97c8232cf319d9c101b07b2f05d215764c38ab30cf9

Observation db9b52e4-04f1-464d-9c3b-854d8b45cc3f · inbound

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation cites this paper.

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.971946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T03:32:23.838961Z digest=sha256:7c3c9fbe2a01e974e70ee4d9646cfb1af675375f19da997101fb206fd61c8b56

Observation 2cc58875-3fa4-4603-9419-b2cabe24bced · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.452297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.452297Z digest=sha256:4a9e0b42701b730ec06d871cf09934f66f366811f9422b6c224a8a228e90c85f

Observation ce6e5120-c663-4fa3-805a-fd0067f1c597 · inbound

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation cites this paper.

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:08:04.544761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:08:04.544761Z digest=sha256:46bec8caf3bfb00cbb7846ab70efbebe947667148018251ce2d007f5e68e5825

Observation e2f04486-eb39-4b01-920a-b8217a575738 · inbound

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems cites this paper.

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:37.741292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:37.741292Z digest=sha256:c5df7f7e08fdd357e289e247ceb3db448530ef9a0cf890331c25090d9baf4bc4

Observation 5e39b400-d8f9-48ee-812f-f0de8a735222 · inbound

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation cites this paper.

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:43:26.122470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:43:26.122470Z digest=sha256:de2ac8674d66fb6052d4dd82d6d0703ab65ed44058d82a0ab05faaff22d44b62

Observation 1e05fdea-1195-4406-989a-5ef9c53d79dd · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:22.277483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:22.277483Z digest=sha256:cbebdbb3d2da17340abdf844f1a37dffda29999bb60a704f62a2f57094395e4b

Observation 8ad53429-5f47-4f33-9b3a-5a55b1bb9063 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.938774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.938774Z digest=sha256:f775c9280294aeab5492be1d05f6a3e8e765ba274951d91ca5e7942da047ed1f