Pith. sign in

Paper Citation Record · LEDGER

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2507.00324.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00324 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:23:53.153699Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:23:45.554040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:26:22.148036Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact4
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec41e0e6-1bdb-490c-9c9d-924a5e4a3f48 · outbound

This paper cites This high quality of synthesized speech and the ability to distribute it through so- cial media platforms are giving rise to manipulated information in the digital ecosystem.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges This high quality of synthesized speech and the ability to distribute it through so- cial media platforms are giving rise to manipulated information in the digital ecosystem

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:59.161684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:45.377493Z digest=sha256:958b8f024cd323fc5a5d23aca67a8ca80fd3ceb44fa8ceae071e6001c024ef08

Observation f6badd89-b359-41b2-abf5-00ce9b129525 · outbound

This paper cites Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:45.554040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:45.554040Z digest=sha256:dde78a0b4c9596816ca3e8d5a02b5eac43e22a5a6a78722dfafac6e7030481b3

Observation 34a76118-16bc-4390-b374-e4e12431aa93 · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:58.763526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:45.820734Z digest=sha256:d4abd98dba3359b42458fc36a99a6461fb6a66b3ffe74b5284e45ce3d119e282

Observation ef2e435e-1bf9-447f-b5bf-7d23e788857f · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:58.124999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:46.517361Z digest=sha256:d5b318d263a8a36e653403bd545b5df689b47db0078a43b6f3dc4e890defbad1

Observation 0cb07594-6495-4010-8f5f-3ddfbfd46961 · outbound

This paper cites It uses ffmpeg to download the best audio available in the W A V format, and resample it 16 kHz.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges It uses ffmpeg to download the best audio available in the W A V format, and resample it 16 kHz

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:58.593590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:45.966902Z digest=sha256:99934b78c554920f3a8f4fefa1432f618de9123248a0c36d76d90c59b2e8bebb

Observation 1b89340b-2c8d-421c-b373-261544e841bf · outbound

This paper cites This step eliminates cross-talk and background speakers.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges This step eliminates cross-talk and background speakers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:58.415398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:46.148015Z digest=sha256:18acb88b3f670dd0711e0bad1a835fb190144e5354d91ba2db9c7777aa913d57

Observation 538b5387-e11c-4005-ad38-1da8854d1d6d · outbound

This paper cites We also experimented with the Google speech recognition api package and other commer- cial tools; however, they generated text with less accuracy and without proper punctuation.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges We also experimented with the Google speech recognition api package and other commer- cial tools; however, they generated text with less accuracy and without proper punctuation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:58.280280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:46.307871Z digest=sha256:c2dabec0f0cb5a39fdb3a974d012913a7b2f1d1f7c9c759f79e8a108afeb176f

Observation 4c6b9b41-3a0c-48ba-b550-de8cd6d9fb70 · outbound

This paper cites The Biden Deepfake Robocall Is Only the Beginning,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges The Biden Deepfake Robocall Is Only the Beginning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.327407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:48.600644Z digest=sha256:aa1332bc2493eae4d05580b42f1a862cc26dee83c4c02c7135ef8a5a5f0a9ac8

Observation a2d27f96-3a9b-4336-98aa-fb4c88df9e0b · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:57.977394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:46.655038Z digest=sha256:546a6ef92b87a9627ac7583b597b2130804d3b1f2cf6a1512016457301d602ba

Observation 4e2b14f0-ca40-4164-8436-31dbb8330c30 · outbound

This paper cites Al- though early models produced robotic-sounding speech despite extensive training data, SSL-based methods significantly im- proved quality through large-scale pre-training.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Al- though early models produced robotic-sounding speech despite extensive training data, SSL-based methods significantly im- proved quality through large-scale pre-training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:57.786761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:46.800698Z digest=sha256:6afdc8b0c3772cbe30082ddcc5fddb0af5239de438b0fdc1f7415b12fe41309f

Observation db3e056c-b005-4bf6-be5d-33f6bd23d504 · outbound

This paper cites We train StyleTTS2 [22] only using this approach.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges We train StyleTTS2 [22] only using this approach

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:57.612401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:46.967101Z digest=sha256:e9edd8dd3adcb9d7d8f358e4bace022dd091ef1423ee4d69641189d9c5b04a64

Observation 68170591-9c61-42c8-965d-3b9d503f8dbd · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:57.454318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:47.155081Z digest=sha256:a910d95cb40740c8fe53b914bd27ae05056abaa7d82d8cbcb6d48fa1d0b43933

Observation d3b137ea-e398-49a9-8790-0dbbb28585d5 · outbound

This paper cites This approach preserved linguistic coherence and improved phoneme alignment, leading to reduced noise and more ac- curate spectrogram generation.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges This approach preserved linguistic coherence and improved phoneme alignment, leading to reduced noise and more ac- curate spectrogram generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:57.305606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:47.266100Z digest=sha256:74d31a4f7dab7f0a4a3609d7bfe75effd11c47e80e5d7d8d39d9f142e9a323f0

Observation c11b0556-b91d-47d2-b848-ecabd713c468 · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:57.171255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:47.416615Z digest=sha256:9f5a864ee145a77d546f25834f3138289f0ba19c6ce40066f500c5ae52fa62ef

Observation 7963db6e-35d6-4ed1-8b50-8ea095bc9efe · outbound

This paper cites For subjective evaluation, we implemented a web-based listening test with 32 unique participants.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges For subjective evaluation, we implemented a web-based listening test with 32 unique participants

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:57.026356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:47.567771Z digest=sha256:538f3331b462f60e6863497c549be5f50e4dbda31b9cfa572edfe812ce88c06f

Observation 7c026615-d60c-4f4e-b76d-e32cb31a84dd · outbound

This paper cites Both datasets are derived from the VCTK corpus, which comprises high-quality speech recordings collected in a controlled laboratory environment.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Both datasets are derived from the VCTK corpus, which comprises high-quality speech recordings collected in a controlled laboratory environment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:58.981915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:45.684201Z digest=sha256:eec5024d4f51141c141af854e087fe1490b98b15bba972ee6ff719ce95b88197

Observation b33136b0-dcd8-4c67-b4de-efbb7d05da97 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges A Survey on Neural Speech Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:47.667605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:47.667605Z digest=sha256:e5c9850e6311361d4b1b2ca3a78bb4b1e206b3fe8c3845c2d7629c11a8295aeb

Observation 8341b7aa-fa69-4608-9424-d0d851b6f797 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:47.831204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:47.831204Z digest=sha256:b92885b13251a8083f75c8f20a71094570cfc1cbca4377b4863891ca7225a9fb

Observation 42a44163-1117-4c28-a7f1-e42f8587f90f · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:47.947919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:47.947919Z digest=sha256:a244a36f7ecc601f9eacec7972e6ee29ce27567299fb267fda38ac4fc65fd0f6

Observation 013f9f70-da0e-40f7-ad16-7ecb4c6027bb · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:48.037703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:48.037703Z digest=sha256:ab8c00e33eab6ba60df8474a7737a08262589ec787f70e2acecbf7bb1ee2f701

Observation b4c32037-b3df-447c-a59e-bf9d1295845a · outbound

This paper cites Zse-vits: A zero-shot expressive voice cloning method based on vits,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Zse-vits: A zero-shot expressive voice cloning method based on vits,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.862210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:48.196907Z digest=sha256:8f8b734946f4147ee1accc91679b8380f1e1755ea2f1a5478b5d29781cdebb7d

Observation f633927c-eb74-431a-9d15-b2a2f1f499a7 · outbound

This paper cites Global Risks Report 2024.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Global Risks Report 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.696718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:48.363816Z digest=sha256:4c567d8bcc94d2a6b08dd2ae4984524f3c6943e0de87ddcf3d0832b926f9b4d0

Observation 0646b28d-b632-4307-8403-ed1575d8c012 · outbound

This paper cites Russian War Report: Hacked news program and deep- fake video spread false Zelenskyy claims,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Russian War Report: Hacked news program and deep- fake video spread false Zelenskyy claims,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.520757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:48.508515Z digest=sha256:0c7da685f98a3d5edc7e6969f90e523562c12561fe435f15bc451f3911a4aad6

Observation e58cadbc-45e2-4815-8c6e-4857dd46ce3a · outbound

This paper cites Sadiq Khan says fake AI audio of him nearly led to serious disorder,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Sadiq Khan says fake AI audio of him nearly led to serious disorder,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.151468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:48.786228Z digest=sha256:15ce12f483d59cd1b7b36959067222644e07ff0e75a787cd11872afeeaf6d517

Observation ec2c84af-dc38-4b29-94ad-3daabec6004d · outbound

This paper cites Asvspoof 2019: A large-scale public database of synthe- sized, converted and replayed speech,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Asvspoof 2019: A large-scale public database of synthe- sized, converted and replayed speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.981769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:49.011538Z digest=sha256:45c026a67fb581d0a90f87a9843cb8e3337c60f7814c429ce63b6b9fe7530b83

Observation 9881999c-8420-449f-ab10-23fe9167081d · outbound

This paper cites ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild,

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:23:54.270034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:49.236797Z digest=sha256:3c1a69bf30c35b6cca09f4472304784dcf40bdba150f2711ad654c9ad84ca2be

Observation 6d8b812b-97c3-4995-af76-0fcc54dedc96 · outbound

This paper cites Asvspoof 2021: accelerating progress in spoofed and deep- fake speech detection,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Asvspoof 2021: accelerating progress in spoofed and deep- fake speech detection,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.809406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:49.402291Z digest=sha256:babea272cd3f45a313efcd003f7f4b0c83fffc40ae2caedb9bb6db12f38fbf3a

Observation 895e6283-bfc3-4fe6-a159-f39d9181dde1 · outbound

This paper cites Asvspoof 5: crowd- sourced speech data, deepfakes, and adversarial attacks at scale,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Asvspoof 5: crowd- sourced speech data, deepfakes, and adversarial attacks at scale,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.597029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:49.475227Z digest=sha256:c5614de29ca948f7b5c2239684b9ac5bcee85fb5310374200af2c096af853ff8

Observation 92954748-fe91-46f9-8a87-5e3657184c02 · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:49.572150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:49.572150Z digest=sha256:3c126e84c4d28a24de2bfe479813c09cc342c5b681b0ba97a702cd6b2764976d

Observation c35826a2-ee87-426b-bcb1-1ba8811c9704 · outbound

This paper cites Is audio spoof detection robust to laundering attacks?.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Is audio spoof detection robust to laundering attacks?

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.425051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:49.650417Z digest=sha256:bf5227a2a033f89f7934f8e8644a5a7a5710a1a1c50f9385d12767511ac2f692

Observation 98d7743b-a51d-4fd3-90cc-2ec3146a1361 · outbound

This paper cites Dfadd: The diffusion and flow- matching based audio deepfake dataset,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Dfadd: The diffusion and flow- matching based audio deepfake dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.291982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:49.755413Z digest=sha256:d97923f08e981054ebf429d9c9fa90fbc333e5e7269b231fe4fd422942f5f75e

Observation 202f6993-73f5-4b3c-8b6d-764782f87ef3 · outbound

This paper cites The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:23:53.920068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:49.860090Z digest=sha256:0ee9fc8b4eb940f838e82282c46418fa41445d06e6d2f4781e9782b22743c9e7

Observation 59d04091-d3ea-4307-94a5-356266d1489a · outbound

This paper cites Mlaad: The multi- language audio anti-spoofing dataset,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Mlaad: The multi- language audio anti-spoofing dataset,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.120571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:50.027320Z digest=sha256:8ca8698cbd82f1cc5566f1fca668bfecad2b985e55e051565bb55d6685b13e51

Observation 016cbdbd-4287-4066-bc7e-543d76fab3e5 · outbound

This paper cites Does audio deepfake detection generalize?.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Does audio deepfake detection generalize?

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:54.952316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:50.161862Z digest=sha256:3e22e727ee66f0ca30de80e21b194695e80d0972a7317a28f496f776ecc73083

Observation cd5d01b6-2b31-4925-a019-4ed5511686c0 · outbound

This paper cites Spoofceleb: Speech deepfake detection and sasv in the wild,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Spoofceleb: Speech deepfake detection and sasv in the wild,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:54.824184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:50.340067Z digest=sha256:b7370a43b12171b4076b76646a84c530bce0281d3b84e43656712e8b1c0a9efc

Observation 285295d0-de4a-4fb1-8bc4-32d9917fa255 · outbound

This paper cites V oxceleb: Large-scale speaker verification in the wild,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges V oxceleb: Large-scale speaker verification in the wild,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:50.513280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:50.513280Z digest=sha256:524b7deb5e3219ba125f47e37b30043722d4daf56e89e9666624de0135fb9d06

Observation 3d92ef0a-b118-4ee8-92f0-503f3c09a367 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style dif- fusion and adversarial training with large speech language mod- els,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Styletts 2: Towards human-level text-to-speech through style dif- fusion and adversarial training with large speech language mod- els,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:50.723911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:50.723911Z digest=sha256:d39431577f32da5c49bdf2dd05144a12ab0ce0ce63b7b57f4b47dedd18f8fc4a

Observation 6edb5495-89ea-4980-8374-716e9984d724 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:50.893642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:50.893642Z digest=sha256:509eabf38e5aa0302cc7875cd6004a5e62bb268355aedb9b6ac40dbd9ae13328

Observation a5447dd1-93e1-46c0-a6c2-e656c2f2e369 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.063898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.063898Z digest=sha256:5da44331f67eb2be43541aba519f15deff2f3f8ed2309739eddfdb7986bb86b2

Observation 298bbffc-0c0e-4256-9e0a-b1091923898c · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:54.634215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:51.260561Z digest=sha256:27b4f4cee46762bb4ed74f7459938aeda2f905e3c8e01e05c5a6b0e3ce8d2b50

Observation be5a0f09-b588-4d3b-be6d-6f1fa0693359 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.398868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.398868Z digest=sha256:5fd11b7c8f098d37c3a1a85a4b2dac7baee7d8f998b8dce54b9144264260b794

Observation 02eeabf4-9992-46e8-b95d-3e9fb890efc7 · outbound

This paper cites SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:23:53.712615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:51.548122Z digest=sha256:c38ea6175b9ccd97f87cdc808ce0dde2e2acea8aec03eeaccbb5f7c3cb7d9fa5

Observation 824d47ee-60fe-4859-b56e-c26291a5af5e · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.703103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.703103Z digest=sha256:f7c35c6f3777c0afc0116fab60aba9f66eb595ab819dac57eb4d665997a1098d

Observation 9ae2daa9-6aba-4670-8a54-0512751c1724 · outbound

This paper cites Cosyvoice 2: Scalable streaming speech synthesis with large language models,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Cosyvoice 2: Scalable streaming speech synthesis with large language models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.927221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.927221Z digest=sha256:5ee2629d0faa6d2518e39f2dd4977186f9ea920dbf59560241e005ce4067f14c

Observation a0716cfd-1be5-4404-b1eb-a91a8fb0ecfa · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.244690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.244690Z digest=sha256:58af93319aee80e22b77cfdaedb3ef2b7bb0216eec3dcfb0a0eef52d07d6f488

Observation 09561bf6-690f-4de4-9bec-0780c6c14941 · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Tacotron: Towards End-to-End Speech Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.397089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.397089Z digest=sha256:b269efb4d95c86fca5b2f1c0474364c831745b3b650c824d2128c22170cb4058

Observation 089e98bd-4114-4964-bf2b-ecf40630e816 · outbound

This paper cites Glow-tts: A genera- tive flow for text-to-speech via monotonic alignment search,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Glow-tts: A genera- tive flow for text-to-speech via monotonic alignment search,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:54.495771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:52.542844Z digest=sha256:aed51dbca2bddd8b1216e9cc8a7e43d8be788e40a6437c124f61af10c7d801f6

Observation 416732da-24e4-4e09-9e70-d2114e6776df · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.752336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.752336Z digest=sha256:8be179c7309396494c24fe81bc29062f3d0ae009b851fb29c6cbda59d58e0bdf

Observation c8af0324-78ea-4ecb-b44b-104148aad9b4 · outbound

This paper cites UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.972604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.972604Z digest=sha256:4a1ef10860cb504b4e5ebe5d884c229ab206342103d784703eab4653269b7e43

Observation 1dd4bc21-92bd-4520-bd50-7f7ad9e468a6 · outbound

This paper cites Deep Learning Based Assessment of Synthetic Speech Naturalness.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Deep Learning Based Assessment of Synthetic Speech Naturalness

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:23:53.453912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:23:53.153699Z digest=sha256:44fe6aa7b416c18ef55f45cbf6f57f9f5d8af551e82081b0143e67d1a84291bd

Observation 94fc393c-e3fe-4e56-996a-d6340c2ed578 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.083979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.083979Z digest=sha256:859b720d22662e0ee294d8f6d51d59ae61d1dca8ea681ea073cf5d75ab694a52

Pith citing papers

Observation f6badd89-b359-41b2-abf5-00ce9b129525 · inbound

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges cites this paper.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:45.554040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:45.554040Z digest=sha256:dde78a0b4c9596816ca3e8d5a02b5eac43e22a5a6a78722dfafac6e7030481b3

Observation aa481225-ba01-4ef9-86ae-5445907a2a84 · inbound

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection cites this paper.

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:26:22.151421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T17:25:27.137423Z digest=sha256:f13f34d2acf959c33e9a726f7b8d66aeb8bc05ce6b96ad69c6a222157d12e3b9