Pith. sign in

Paper Citation Record · LEDGER

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese

As of 20 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.11200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11200 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:02:45.995503Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T12:03:11.243693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:21:09.992559Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact4
  • verified fuzzy26
  • unresolved17
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d12e0eee-0575-4604-92f3-acec6972d62a · outbound

This paper cites The kendall rank correlation coefficient.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The kendall rank correlation coefficient

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.916722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.452290Z digest=sha256:589504a32399c00863b0d0373a548853499e501f7ab51952500647424b47086b

Observation 78e42f87-1c18-4a06-bbc6-b9bf84f4f46d · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.460507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.460507Z digest=sha256:fad3dca104225440ec8a5aafc5d9eb55cee62a342a96ee136d71870df4984f3c

Observation 49ab2a14-7b56-4a04-961c-c6bec3e765b2 · outbound

This paper cites The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.889436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.474393Z digest=sha256:fe011e92ba587cdab0eb4ff1a138d30720539dc5ca2aaa9e5fe72d0879ac9762

Observation 27821350-9168-49bc-855e-bdb93b82c704 · outbound

This paper cites Generalized linear mixed models: a practical guide for ecology and evolution.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Generalized linear mixed models: a practical guide for ecology and evolution

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.854107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.485384Z digest=sha256:4d445741f61685fc2a6060c24c4d50ac25dadb82db42b44205d3b6b86c29ad7f

Observation 2feb102f-9eb0-4a89-9a40-03be0124b9db · outbound

This paper cites Why we should report the details in subjective evaluation of tts more rigorously.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Why we should report the details in subjective evaluation of tts more rigorously

Reference 5

Resolution
verified exact
doi, observed 2026-08-15T21:02:46.147655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.499141Z digest=sha256:81718fcb50e3494ce81e63e9cc60dc22e0eea4ce5857c73a8984dab0c56ae6ff

Observation 34a5caeb-448c-4616-b27d-41c5bec6b7c2 · outbound

This paper cites Qwen2-Audio Technical Report.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.507556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.507556Z digest=sha256:4f7727a46bf86f7446402768a97816dd9ba22d8a85a6c2ab76504b59374bb20b

Observation 54564007-9d02-4775-8abd-164db0a99811 · outbound

This paper cites an unresolved cited work.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.519126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.519126Z digest=sha256:67092c03cd233a76cddf06eabb945e4e6b4515abd3fc953320efaba161a936ef

Observation bd94a023-0738-4b32-85ff-82892eab48c6 · outbound

This paper cites Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.527668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.527668Z digest=sha256:5a2102b5674021a15ef96640b972363730d4300aabaf8748bb0f348e37e3da53

Observation c987edd5-03fb-4ff8-bef5-4dd278e08814 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.536218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.536218Z digest=sha256:6a077de30a6849b90ddeec020618f47362943b453e6f09afadcf49f1435eb75a

Observation 939b2ce8-a515-4760-b107-dd68c9163cea · outbound

This paper cites Assessing the impact of contextual framing on subjective tts quality.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Assessing the impact of contextual framing on subjective tts quality

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.824182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.544926Z digest=sha256:b413d21aedcc5c1fb4ccea8f1cbdc9bed10ed0936b4403a9d7afa5d65d791fa6

Observation 075d394f-af28-4007-8479-c4a61a262681 · outbound

This paper cites The turing test: the first 50 years.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The turing test: the first 50 years

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.789140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.556597Z digest=sha256:5ffa043dcd0fbdef3fa7d76648b02540fd9de99faff443b1d9ddcf8ec4a1f20d

Observation 0bcd6c3c-53d4-4c3a-ae1a-db140529cb68 · outbound

This paper cites Analysis of speaker similarity in the statistical speech synthesis systems using a hybrid approach.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Analysis of speaker similarity in the statistical speech synthesis systems using a hybrid approach

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.754777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.566322Z digest=sha256:bfb4142318c14e3a80b4a8e6cbf7bb6591c6b495751ada9d38fa7c729bf42470

Observation 3f9bdd01-ed8a-49c2-83cc-865a0096be17 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.572675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.572675Z digest=sha256:5ba7f5859206e6496bcfe4bc355ca77697ce54044a60aff674df563f69a9e7b3

Observation 838d049d-4211-44f4-8e20-80ca2dd798ed · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Lora: Low-rank adaptation of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.715922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.582227Z digest=sha256:ae9c92c68df05c0f26f689c05a2010cc20a1f9c87d8b3e24da55a7c341c9e671

Observation 0873eeb2-a2ab-40a1-97d6-4a8611358ce0 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.591082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.591082Z digest=sha256:66008b406c3c708fa9346eeb662a9465f4536268d599801097fc0242ab7e3cbf

Observation a2c57a22-089a-4124-a412-99066bacda59 · outbound

This paper cites Mm algorithms for generalized bradley-terry models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Mm algorithms for generalized bradley-terry models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.599714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.599714Z digest=sha256:5063f4884e126a35acc771b120319c30c00926dbaf000647029e6b75ebb777cf

Observation fc5ef76a-f15e-4309-ab88-086bec1649de · outbound

This paper cites GPT-4o System Card.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.608954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.608954Z digest=sha256:f0b35258a8b8a22f52b2d85c63c1b33d6e529a9ec3d4609681eafcb9e502ed76

Observation c87f2b6c-f133-4c6f-9ddb-07c0a1ee6fa1 · outbound

This paper cites Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems, 2015.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems, 2015

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.634139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.617482Z digest=sha256:16e3dbb8214c0bf3d29700148558e9d2df554c593b2d16047d2ab9478a01f8cd

Observation 71107992-3aec-49a1-a264-a01315bd3f0b · outbound

This paper cites Subjective evaluation of speech quality with a crowd- sourcing approach, 2018.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Subjective evaluation of speech quality with a crowd- sourcing approach, 2018

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.609343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.628694Z digest=sha256:e88fc462c9fa55773522e642d3ff10fff38c2e3963660107c12e3bd99c934d47

Observation 8918ce2d-48dc-44f8-83a0-fb5157157b10 · outbound

This paper cites Compact Neural TTS Voices for Accessibility.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Compact Neural TTS Voices for Accessibility

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:02:46.619507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.639186Z digest=sha256:6582aec469c99ae0c751690a55b8f25fcc0a65cce84318bbbc0c35905fd30720

Observation 4fe8b601-769d-4554-adff-150867f3c357 · outbound

This paper cites Stuck in the mos pit: A critical analysis of mos test methodology in tts evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Stuck in the mos pit: A critical analysis of mos test methodology in tts evaluation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.569430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.671587Z digest=sha256:8854629ab72a0a3001c0c3f49bfee8a493dcd7a145c679433931d8ff3cad15ce

Observation 78dc6c6c-af68-4195-ae55-065e6914a91e · outbound

This paper cites Issues in chinese prosody: conceptual foundations of a linguistically-motivated text-to-speech system for mandarin.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Issues in chinese prosody: conceptual foundations of a linguistically-motivated text-to-speech system for mandarin

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.540103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.679355Z digest=sha256:fe018957fe6fa936e878711f1e0e9a3c89edf25eeef8ec895c66c800f9481ad5

Observation 1ddbb4e4-8785-4c2b-9de7-9bd1c2c331a4 · outbound

This paper cites The limits of the mean opinion score for speech synthesis evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The limits of the mean opinion score for speech synthesis evaluation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.496958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.686752Z digest=sha256:7caf9ba8ea7cd733eaefec194733952018e80525dafdc81591734cdbd9a1c61b

Observation 25e06993-00be-4795-9ff1-7e7478a64e9e · outbound

This paper cites StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.723124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.723124Z digest=sha256:fc44d03e37d6c2c27e6c849bc90d8a78e36a4b626cb0dadf23b91408cd6c99fb

Observation ee85daee-7f91-4448-bb93-d3020af5c143 · outbound

This paper cites Hyper-realistic, multi-emotion generative speech model speech-01.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Hyper-realistic, multi-emotion generative speech model speech-01

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.462948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.736670Z digest=sha256:db57caa45fdebb5dcfc7b5648372ad57ef00855591af7f679b9e1a45c8627fa2

Observation 23f73e61-8485-4b1d-b7a4-b63c8ebd9b16 · outbound

This paper cites Speech Quality Assessment in Crowdsourcing: Comparison Category Rating Method.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Speech Quality Assessment in Crowdsourcing: Comparison Category Rating Method

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:02:46.404802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.765487Z digest=sha256:3a25552ce7071879552e1879631a635cb712ccbf6d5fc2fc7e6b03c1327be126

Observation 653dc60e-7403-48af-a90e-c356fc618f18 · outbound

This paper cites The blizzard challenge.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The blizzard challenge

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.427943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.773130Z digest=sha256:786555aa38a35635f75829b4e7135375e69457942cc6666900e36ecf312b22dc

Observation 5ecee893-2d36-4eca-a73a-a04941c003c4 · outbound

This paper cites Dnsmos p.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Dnsmos p

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.401200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.796723Z digest=sha256:5014bf10b1dc8811d006184a44ba943537104ae607fa5522c5369fd851fc526c

Observation a8de6606-9795-4a78-bbc8-7041e0f01c01 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.817444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.817444Z digest=sha256:28466cbdf8943338043d00b8ecf158c677d81256afea4c703eb815b092312632

Observation 7ec24e69-2d0f-447d-9579-242ec86fb6cf · outbound

This paper cites Mean opinion score (mos) revisited: methods and applications, limitations and alternatives.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Mean opinion score (mos) revisited: methods and applications, limitations and alternatives

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.357552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.827440Z digest=sha256:1b7b741fa7442cb5ca6a899dfb01c1222790d62cd1b683dcca6beadf4683f42b

Observation 2d921818-f727-435a-b522-005c737f94ca · outbound

This paper cites Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.838352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.838352Z digest=sha256:36378b541a165d2dd1840de4246b047ab9b3805cc8829fd29cb108bed0b89399

Observation eca95645-6b86-4f9c-8307-7e66d4703efa · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.850786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.850786Z digest=sha256:8518f601cde4c9658f0f3c47260cedbcad214a57ebc707b89f4d7bad5fda15ac

Observation ed67bc23-3e89-4bb0-bb63-f9fcfed3ee86 · outbound

This paper cites Contextual interactive evaluation of tts models in dialogue systems.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Contextual interactive evaluation of tts models in dialogue systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.320286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.859506Z digest=sha256:6e8036c364d45f172ac9c773b96949b96bd5092c3824ee9bc5020933da973cb0

Observation 320ec6cd-47ca-4755-b7f4-245507619faa · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.866945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.866945Z digest=sha256:aae424144b0f98f7a176d20a6c736cc42a0b7ca146d5c0949e97df1caf5f0e59

Observation dfbe0640-77c6-4aff-81a3-0eb3fbfc9e74 · outbound

This paper cites Bilingual and code- switching tts enhanced with denoising diffusion model and gan.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Bilingual and code- switching tts enhanced with denoising diffusion model and gan

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.271146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.875017Z digest=sha256:1c67a39c92d5b051a09af15b319ed3f5627e91b7a33dbf8b18005c1cc2a7a793

Observation 7d5914c6-f227-43b5-876d-6e43de4d5959 · outbound

This paper cites Talk2care: An llm-based voice assistant for communication between healthcare providers and older adults.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Talk2care: An llm-based voice assistant for communication between healthcare providers and older adults

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.883992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.883992Z digest=sha256:57a20ab8a40eefc13a337f4bb72db5d185699349a3996a55f8d873d6cfc7ed7e

Observation 0e0df59f-fc67-4124-91d2-ab4338032a0c · outbound

This paper cites Dialog modeling in audiobook synthesis.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Dialog modeling in audiobook synthesis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.228324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.896901Z digest=sha256:45ef0968b22042c49e01026afcfdc18bf3da977ef350fd5a3e07b081d76fe737

Observation 730a5761-7d6e-48b7-b7f6-fc5a6f450973 · outbound

This paper cites orders of magnitude larger.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese orders of magnitude larger

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.196380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.906044Z digest=sha256:83a87bf108e8081e3c5cfa413a0f48866f2b2d7b5ef962485092e7546e11eecd

Observation cdbc3f11-774e-4e02-b821-670445f8fc91 · outbound

This paper cites Pure machine voice.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Pure machine voice

Reference 41

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T21:02:47.161308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.913721Z digest=sha256:217edc8dfc510922c57daeac528ac65b5c77d482832cd53ca1ed32bb8f198d87

Observation a6cc27a3-e201-445b-8c9d-6e51f2a837f3 · outbound

This paper cites The i m i t a t i o n of human speech is too forced.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The i m i t a t i o n of human speech is too forced

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.135018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.927549Z digest=sha256:376f189d422392af9ea45173b5d94550e71c5b7be5b23c05285900cbc226a7de

Observation 64a8d3f5-6982-4628-b4c9-dae1e669ed74 · outbound

This paper cites O b v i o u s l y a machine tone - doesn ’ t sound like a real person.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese O b v i o u s l y a machine tone - doesn ’ t sound like a real person

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.098970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.934317Z digest=sha256:ae2641b5acb12c78d19ef11e0834205695d6676c19aa18f234e55459b71c59dd

Observation da151159-e479-4e0f-999d-208f638af4c1 · outbound

This paper cites Sounds like a late - night radio host.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Sounds like a late - night radio host

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.071983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.948407Z digest=sha256:0db08db3c172e69ae595b08b7fb78ce7f487c0cfd599ae6937f430c5541017e1

Observation ac974039-4221-4c92-afeb-5000f6d26a52 · outbound

This paper cites Many thins.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Many thins

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.047847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.960734Z digest=sha256:6c6d27d422d2041ab9144c35b01515b89aa9c08df70f899f03a2e929ce297e7b

Observation cc2d41db-d096-4f5e-b35b-038e7cb39e49 · outbound

This paper cites an unresolved cited work.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:02:46.992756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.971102Z digest=sha256:c2b8a4b033d1afa52cd618b450a08f6c4ba29e4c23b379d9128e7bc4763ce2c6

Observation 69e6cc48-396f-48e5-ae18-3bec302d2417 · outbound

This paper cites go away.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese go away

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:46.948182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.978046Z digest=sha256:3da3f92fd12e5cf2090b2759739785930ecaffa31a18c58d10ff943e8cc7d367

Observation 730e009a-eede-46dd-b78a-3a67896de368 · outbound

This paper cites angry ,.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese angry ,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:46.921239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.995503Z digest=sha256:e05d3b6d22ca3d848fd1b98db588b03f5fb15111edb00242547e5beb2c0ee3e3

Observation 45f9dab7-753f-4a68-a415-96c670f0e988 · outbound

This paper cites doi: 10.21437/Blizzard.2023-1.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese doi: 10.21437/Blizzard.2023-1

Reference 2023

Resolution
verified exact
doi, observed 2026-08-15T21:02:46.092266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:02:45.789007Z digest=sha256:b03c8824e095e9ff3b3c413c9cda9451a9e6d32d0b4f28a8a5e38df0a13fefd1

Observation 1aed6274-3448-4c8d-afe2-64eb5efbe64b · outbound

This paper cites URL https://www.sciencedirect.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese URL https://www.sciencedirect

Reference 2308

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.696157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.696157Z digest=sha256:08f5efbcca5079d8d1472e7fcf1d299f1a8c8a4daa7ce6e5a84ecce1fdd8ab38

Pith citing papers

Observation d1e43662-3b5f-4805-b13a-e882f93a10a2 · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.999306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:844efb8bb75b59f9e306399f8f32f19d633729c2e44d7545d3531a220d40555c