Pith. sign in

Paper Citation Record · LEDGER

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

As of 21 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2505.04113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04113 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:42:53.805578Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T03:50:26.873406Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1a7df6e4-fbf8-4191-8a8a-c79b921f0ec1 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.382881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.382881Z digest=sha256:f07b94c5a56fdde3438817f06cf06f21658ea02be9d5be02ac0e82bacf3cd9a9

Observation 4e61aeb9-fa24-4000-8acd-1f5de6c1e587 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.429120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.429120Z digest=sha256:7cf5126958100da6d987a5e1ffbaf3a81df46b856dc71cdfb1b37982fc60c4ca

Observation 0fec1e81-ec53-4289-8fb0-9aeafa1c6848 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Common Voice: A Massively-Multilingual Speech Corpus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.435184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.435184Z digest=sha256:60cd48134e5cb2f42714a0f28630b44466311013ec7eb296aba247fe59922fc4

Observation c054b9b9-5f4f-40f6-8551-940fe0138671 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.440665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.440665Z digest=sha256:5246a407ee96e6cdfc4ff04d82e3e674e6e2dde6dd6bf13c6c43d3ef56a73c43

Observation 4a136a66-6ecf-46cf-aa10-4af35ee3b71c · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.090896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.446506Z digest=sha256:0ab4cd33125d3cc66075f7db5170d1a27452fd6b66115f0be6e1ca3cbb961cec

Observation 8a4fdccd-eb9b-4ad5-81a9-c3a33c7f376e · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment SoundStorm: Efficient Parallel Audio Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.451504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.451504Z digest=sha256:17f90c793e019f541cda28f35aef5f2334a25eeb0f2eb9a037b99a3f28804625

Observation 8d8b4098-e038-4e14-9b6a-916c39232cf1 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.074759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.457377Z digest=sha256:5e3e0fdacc705c712509de71b9cbc57c73498154bac548546afa4096307c0050

Observation 08898a01-1af3-4201-a8ae-8b94ada6e9b9 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.462636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.462636Z digest=sha256:bbe83f7b6ba2296d1192e0cb95834d7867362cbbf58b7f1bbe5a64db4a9328f7

Observation 550ed16d-b6da-440e-87b5-8eb46b30387d · outbound

This paper cites DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.468382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.468382Z digest=sha256:14a33208c821436a2e5bb0edb4ccbcadd455f3f9ce7f891bea908deafb3f40aa

Observation 3f7da104-808d-4685-a7f4-e6cbb78ac27c · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.473616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.473616Z digest=sha256:7c7c98d3dd06133e64f66333ee97b5ea51fc7e955fe4513e7d046e08cd262f82

Observation 0b384649-ebfd-481b-aa21-a76293ff8383 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.478975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.478975Z digest=sha256:9f91cb2ea21cfae0577e66ffe051d433da4e72c44dd8fee755b3696ddcec2fbf

Observation 9c450d1e-8039-4c40-a742-23278d67cec3 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.047622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.484988Z digest=sha256:6e28c6f29da3ada68fe4ae03bac2e398a02f948c7b947f80cf35afda2609699b

Observation 4193e79f-63bf-4d06-b964-160b40e1f1d4 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.489615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.489615Z digest=sha256:87c33d7d9b5bcdcbeeaa53413f620df0a5ab24a0790cd66a81b828d8179acb66

Observation adaeb182-26eb-4d52-8343-0b7745adbe83 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.020747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.508501Z digest=sha256:3149eff589835cdfc6c8c49e500a753da7d9803eadd79dc8f3c90e8dd241a3fe

Observation 227841fa-f122-4f82-a3c6-48375cb356ea · outbound

This paper cites DeepSeek-V3 Technical Report.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.515086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.515086Z digest=sha256:7cbe21a259b836a4d5fe5429e09ec7e1d3fe38066e5f44088891a581156a4edc

Observation bc12a36a-9d15-4775-ba5b-8aec19a18020 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.004563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.519735Z digest=sha256:b13e334067af8d15c86292bbf1c7e97f9173f12214bfd9c205359a447e1a2de5

Observation 4b9d0206-f088-47d9-b4de-4fc0bb4b75d1 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.525148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.525148Z digest=sha256:effed821239ed5be37c3fc9de9709b50336066f41f0b05d9a2ccd234bf93266d

Observation b974ab77-f39f-4698-aa0a-3819e789bd47 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.531276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.531276Z digest=sha256:9c21d38b101a836a66d60610d6f56c7338fef4b74b2fdedd8a973e43faa1ebdb

Observation 548bbc03-0ef8-4a2a-8ffa-0f6d91197ae8 · outbound

This paper cites The Llama 3 Herd of Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.536759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.536759Z digest=sha256:7423478163f2e337cf56778f02c3c730797d8af00bcffbf4ea9891fea7d8f6d3

Observation cabe3b75-9f3d-4e20-b3e4-bf3f270eb961 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.989490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.542162Z digest=sha256:614396ee199be766fb4a69766c1c8b8e5e70c32e90c73ca7371ba6990df3556b

Observation 8dd4a8ed-5495-406d-a597-ff41b1f19ef2 · outbound

This paper cites TLDR: Token-Level Detective Reward Model for Large Vision Language Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment TLDR: Token-Level Detective Reward Model for Large Vision Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.547349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.547349Z digest=sha256:fba4769e5e4577c9650a1627b251ac2d1ff4d404a529c52ea30eedd35c23e0ca

Observation b4a844c4-906c-45a7-835b-d867ddb0714e · outbound

This paper cites Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.552909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.552909Z digest=sha256:0272721ebf47338d6b93e8f53c328e65e02f67511a038b9f921ad2c80d0dcd3d

Observation c374b99a-d696-4c37-92a7-28badd20c2da · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.973191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.559248Z digest=sha256:0f41c63b812dd0c4e64f944167421e83ed091bec3866234da9c898c4890c4bad

Observation f7d061ae-9426-4c05-9c45-c5d549e87bed · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.957150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.564993Z digest=sha256:fbfcfc954006fe1416dd2bed13b50e07d06e6a9940ce14a1cf094d2936cafe3f

Observation 10f9e75e-9f22-4a58-99ac-41179a7d52cd · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.570692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.570692Z digest=sha256:c56b045d9c271d87d431305dc6ecbd44edd1311fe746335aef7e5ecc580c9525

Observation fd9dc5b2-1d60-4564-a35a-0034003851e0 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.941613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.576247Z digest=sha256:525fb9b114f7457cd42cc6cd2b1ad3a355b31035adfa69b22a87e36952ebbc60

Observation f3450993-6148-4fad-91e6-89f8786619ea · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.925550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.581917Z digest=sha256:fa6d844c124cd9c48e862f250b326ae79a37c7671493a2576b93db07ca3e2dad

Observation 400dfd54-299e-4dba-825a-95e6c4df6a19 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.587276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.587276Z digest=sha256:9636adad1a6c17d4ba6136c0e70579b050923645a44c6f83361b52244a684852

Observation 089ac67b-246b-4415-813f-8e7d7cc319da · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.592249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.592249Z digest=sha256:ab7e353826aebcb48e4ba48233bebf157267af689bd7dd4004023a279e21a134

Observation f49712de-77ee-4112-ba43-b53ef2517a5e · outbound

This paper cites Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.597794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.597794Z digest=sha256:0b2ca5b92a706c157c3983ad8ea7a8346c771e76b6ad74a651eda28ccb239401

Observation d3b2ccae-7020-4ddd-acfd-201b60c0d5ef · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.909712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.602750Z digest=sha256:bb3dd32d773adb621a17bbbaf089113e3deb55fb0283ed7c9502469ee5b5e4fe

Observation e9c17c1b-7f1f-4802-a957-c16d7ebf814f · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.893366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.607183Z digest=sha256:db9ada98ff50b2dd1c22987b99c5ffa43a71927e34ffe5f05d44d3115043c875

Observation 03548c78-377f-4458-afc5-2db73b8d1b13 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.876839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.612006Z digest=sha256:2a1e376b48015554880959199daa0c3f50d11f312d331293c46bfd24a5010df8

Observation bdb722d8-41d7-492d-acb8-b1e44dd4a63c · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.859191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.616509Z digest=sha256:fc5be4eb42458a8f1505b8c7ae022b78ffb265fe86821c2a569008bd57d73f62

Observation e2abf169-c2ff-415a-a94e-9440f8ff8bed · outbound

This paper cites Overview of the Amphion Toolkit (v0.2).

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Overview of the Amphion Toolkit (v0.2)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.621398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.621398Z digest=sha256:edb4b365967dcf60951b35d83db5d281de3126eaaf9d5a6dee9fe231ac80bc70

Observation c4d66328-e242-42d8-bfc4-c8d92eee597e · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.842209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.627729Z digest=sha256:04c3babc5bb22bcf999c761d1a8fcbc03948a879ca4ba483031f75bd6fc06c42

Observation feb0c688-75ec-4004-9c61-584a25c57995 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.823771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.633521Z digest=sha256:913b289254a941357211204e9e2881328ad0ca1fb16a91ae714b6d22dfc24a4f

Observation a2782b34-7f3c-4f5c-aaf6-ea0c4c459a6e · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.807908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.639014Z digest=sha256:d257b13e57472aea555d4599bf3d3e801125a6824a491861332308e9bf363b11

Observation 90679bdc-6143-4dc5-863f-534d97529010 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.791659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.644218Z digest=sha256:59436696970656d178915a6525a39dad1feacb3c78e7245af9093fd7733e5fdd

Observation 2e67078e-9dbf-48ae-b7fc-11c2a8cf29d0 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.775959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.649412Z digest=sha256:c9b9916b8e8f0eca33141e894551b568bddf1c99484ad879d365963d6dfa9f44

Observation fe585f51-30b8-4557-ac6b-81cd3419d739 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.758884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.654555Z digest=sha256:6e3b5dcd2fa06704ba7b88c35f417718c1eb7dc1d663bf55f1375de9406c5699

Observation 39689887-aefa-408b-9359-f2cfac8868eb · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.659510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.659510Z digest=sha256:d73e0a9d5ca7a24cb8d0620dd7063b51c512848c70a70ab5c583b693e9b72908

Observation c96d2b90-d07b-408b-80bd-4d0e5ec3e0af · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.664978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.664978Z digest=sha256:ac6455265ced4f647044c6df41ce748b0dca041adec6d2f614c5b22b629c2669

Observation 479152ae-c86a-4608-990d-c04855847f0b · outbound

This paper cites VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.670578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.670578Z digest=sha256:6a964f036d2ef3274275fcf72cf2748d8e02817a96df84ee2217f99560002d6d

Observation 440dfa31-0be3-409f-b196-f0a88836e22b · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.677136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.677136Z digest=sha256:80cd15add59bf97381c5564d5c0a50fe3c77d2a4fad1c9c034b8b85d6c545e12

Observation e7219b7a-c1e8-4a69-bf7a-79d6ed03e093 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Manning, Stefano Ermon, and Chelsea Finn

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.682470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.682470Z digest=sha256:e0a343ea82a2973edff86caf3bb2c899832f505e669edfdb8b0c85d0ee6ae92f

Observation a4a24be6-86ce-4818-bf2e-2667107f0043 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.698960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.687405Z digest=sha256:5187eb640f865c6b670039a0d88d8f8522a6d755ce4b97194fc5fd5cbf4c8558

Observation 7b849058-daaf-4820-93e2-2be380a934c7 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.680850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.692676Z digest=sha256:464f19dbb80743fd08f674172354cfc2b8d7baac073bfa4740ed8664f51152b6

Observation e6a42e10-03c9-4faf-8a9e-96841871c8e9 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.663032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.697796Z digest=sha256:6738e9605c6e984cfbd44a55e806bb034071650bd056964ea6988be0444f2f0d

Observation 4778406c-3704-4c23-a49d-c691c86f6bb8 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.647101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.702994Z digest=sha256:7faabb52b6877e1e63092ec440522d4c21dad96ae1ed807ac49a375f0c330e59

Observation ac7cc6b3-b93d-415c-ae03-2e4ea62851af · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.631142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.708270Z digest=sha256:8be55013fb6d2e212277bdaaed0b50200f411de06641ccf9194ee170578afd7c

Observation 1f5076a4-30fc-41fd-9eb9-160db611f4c9 · outbound

This paper cites Preference Alignment Improves Language Model-Based TTS.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Preference Alignment Improves Language Model-Based TTS

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.713408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.713408Z digest=sha256:630e8635876d7b2a2a46c2efa4d42a01278d7d88eedfe7618fc79a878666a46c

Observation 5472b71c-30cc-4302-8aa7-a70f2d34de2a · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.615562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.718479Z digest=sha256:4f1c4ce93c1eb68770a6f266b288e3adbaa5ba44148c1e1b65db468db91fe4d0

Observation f2993fc9-7472-4e96-87e7-c7adc9d5e8b2 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.723565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.723565Z digest=sha256:18dcd3e1fa1b13675eedd91d1226f2e587083a30f130593abd0ee24ae58a767e

Observation 07b4463b-47b6-46e3-89e5-055217e8dce1 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.599497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.729052Z digest=sha256:6bcc17c1d8c8e8162651cf7394007c8d5cccb9686889e30ae7bf2057ebadb46e

Observation 448c2c86-a5cb-4783-bbc0-b612765e4522 · outbound

This paper cites Metis: A Foundation Speech Generation Model with Masked Generative Pre-training.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.733849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.733849Z digest=sha256:a2e647413a74fca348615ca82485c6993983116cf46f749bc3b29519597182da

Observation 8737ed2e-2e1e-4492-bac2-4446a732a009 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.739056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.739056Z digest=sha256:9a400dd85b8bd1ae9b3dafe5ccf36e8f110614f8582554c63372e2a0f5b799a0

Observation 2fdcfe42-5cb4-448a-961c-a2f136e1253a · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.573360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.744411Z digest=sha256:eb7da005fc0c01ac176b40ec55f32a99886704a9a1d35c296577c0bc111c3eee

Observation a01d4a64-ce51-4701-ba2b-85d3e5b7c7bc · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.557479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.749155Z digest=sha256:0ebab33734250140e4784b010f62a247b8bdc5b21c38a944e4d5ef6a6892af24

Observation eceea7f9-6814-4320-ad10-83d7f16dc911 · outbound

This paper cites Qwen2 Technical Report.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Qwen2 Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.754630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.754630Z digest=sha256:8e7fbc6588f60e198bb4d80c5bcc6a6834683ef37861948bc91c5e7cdfe2979f

Observation a2e45ab6-71ea-43e4-834e-d74a9a321df3 · outbound

This paper cites Qwen2.5 Technical Report.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Qwen2.5 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.759449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.759449Z digest=sha256:4641857c1d79a2b28ac880c0b54ad4be2a63d5d610f6caa5e8bd0359c555a87b

Observation 2b37bf12-725d-486b-9c8b-11c6521e5cb4 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.764204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.764204Z digest=sha256:eb0bca8b2e16b148f8a799c8d5093fb9a5de8dd584e69e7308167a87ab9d6973

Observation 065e654a-1738-4839-844a-b7ec9c4832b8 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.768989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.768989Z digest=sha256:478c1ba4e82429b3885d842f19e3d6efba04713bd6e766ee3be8ee5fc222fbe9

Observation 6fc6d839-8263-4e84-a789-1b169e45c526 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.531673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.774266Z digest=sha256:0a31c575d89ad812f84db9bd4c13d4c24cca4e303e4b86e3ee1b62b77517066f

Observation 1bcea1b1-1a04-429e-a40c-5ebef4a0ce7f · outbound

This paper cites Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.779240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.779240Z digest=sha256:f3893008c20cca1239847197110f9926ba5efb95a48df6cb6de42a414bfe3578

Observation 2186a10a-f63c-4d32-8698-d5870c802a3d · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.515708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.784486Z digest=sha256:9fc354caf226bc542b1b866c49c607ca92bb86cc77bf7432f31124cd4be7e058

Observation c79f1747-c83a-41f8-b5a5-7e6cb985cebc · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.499871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.789633Z digest=sha256:1578b8358246d1e72675cacc67a25d16b74171805d6517a347ccefb417948ff2

Observation b80548c6-48ef-4bf0-84f1-aaadf282322c · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.483979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.795052Z digest=sha256:dd187ff556ec9dea7fa98bd6daa7d09c27830b4c306c09f47ec2dc6f67f48d00

Observation f78a5245-eab7-4d94-bb44-e49322bd21d1 · outbound

This paper cites online" 'onlinestring :=.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment online" 'onlinestring :=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.800202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.800202Z digest=sha256:13c82b6e72f59f2ca5d16b8655a57121a86f721cc5e832f507d63567201b0ec8

Observation 70d85a7e-1406-4e45-af06-e32a602b289d · outbound

This paper cites write newline.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment write newline

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.805578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.805578Z digest=sha256:df6ea7c3dd1ccd6c59086d4a9f11fccbcbcc08de26015a51d376e142765e9ae2

Pith citing papers

Observation b719acc5-8566-4fb8-a1e6-94edf174f7f8 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.542666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:cd2af3792b12fb63419d6b28f941e1335de70383b8ce6e2f668446ed7283606f

Observation d3e9bbc8-f5f1-4916-948d-b4831abe5629 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

Reference 237

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.266206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:865ff5164231ff0098bf1cfcc284006d25e8e796eddcd22f0e359afac534092a