Pith. sign in

Paper Citation Record · LEDGER

Deepfake Detection of Singing Voices With Whisper Encodings

As of 15 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2501.18919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18919 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:01:29.676332Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7884cc9-152b-401b-85ea-c26c7b257bf8 · outbound

This paper cites ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale.

Deepfake Detection of Singing Voices With Whisper Encodings ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T22:01:29.624531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:01:29.624531Z digest=sha256:98b7dd0c4accadcaf3711fc436fb735cda22878886f4da4c2ecf6e1f28fe89f7

Observation 17e5790d-b201-40ae-aa4b-df60e846e6de · outbound

This paper cites Vulnerability issues in automatic speaker verification (ASV) systems,.

Deepfake Detection of Singing Voices With Whisper Encodings Vulnerability issues in automatic speaker verification (ASV) systems,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.997777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.628731Z digest=sha256:e01ec1481c7af13f48260a08390883fe4f441e0433c6e95193422b2fc88b0f40

Observation 6f0203a1-c935-4c7e-8830-3f8b620a150a · outbound

This paper cites Visinger: Variational inference with adversarial learning for end-to-end singing voice synthesis,.

Deepfake Detection of Singing Voices With Whisper Encodings Visinger: Variational inference with adversarial learning for end-to-end singing voice synthesis,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.987883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.632041Z digest=sha256:46d4eeda2d07646e706ba282170595466719e898306074a4a1a09e03a511bba1

Observation 0894bf25-0541-4c78-890f-6c3a5e419a5d · outbound

This paper cites Diffsinger: Singing voice synthesis via shallow diffusion mechanism,.

Deepfake Detection of Singing Voices With Whisper Encodings Diffsinger: Singing voice synthesis via shallow diffusion mechanism,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.977291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.635653Z digest=sha256:980238f63dc8978a715661cdc715f1c5f0ba6a80a61bd369de4209b887961181

Observation cc59e713-0837-4127-887d-d165d73241cc · outbound

This paper cites Midi-voice: Expressive zero-shot singing voice synthesis via midi-driven priors,.

Deepfake Detection of Singing Voices With Whisper Encodings Midi-voice: Expressive zero-shot singing voice synthesis via midi-driven priors,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.968300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.639190Z digest=sha256:701907354800f469da5c1f98a8f7bfc6187c562b099ebe6f1e0186f896878c13

Observation 3e5aa955-04b0-4e02-8120-77dbf1698e28 · outbound

This paper cites Sintechsvs: A singing technique controllable singing voice synthesis system,.

Deepfake Detection of Singing Voices With Whisper Encodings Sintechsvs: A singing technique controllable singing voice synthesis system,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.959180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.642435Z digest=sha256:03480fd2302e11d24aa03dbd72c7ad6f4715590aff7d5c107c09e47ff414189c

Observation bfad7d8b-22f3-4776-87d5-c8bd9e6fd41f · outbound

This paper cites Singfake: Singing voice deepfake detection,.

Deepfake Detection of Singing Voices With Whisper Encodings Singfake: Singing voice deepfake detection,

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-09T22:01:29.859370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.645756Z digest=sha256:7aed963ec0d1a3433375c8408c55e67392071c523ae4638914cfe415a9f9c3a2

Observation 13a450de-71bd-4c75-8848-df4c53b6df59 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Deepfake Detection of Singing Voices With Whisper Encodings Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T22:01:29.648510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:01:29.648510Z digest=sha256:6753d5577c8eaba65425d929fc3046c8fa6c78c33252d603d9488987c2e006a5

Observation eaa66595-6426-46d5-b8db-de5e8d06bdc5 · outbound

This paper cites Whisper-AT: Noise- robust automatic speech recognizers are also strong general audio event taggers,.

Deepfake Detection of Singing Voices With Whisper Encodings Whisper-AT: Noise- robust automatic speech recognizers are also strong general audio event taggers,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.943528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.651539Z digest=sha256:e4a538fcef8e35ef763ad0a2f4213008e09d51dfac5e1cf4c008b1370d554de4

Observation 2627ff73-870c-4990-97fb-dfb158136340 · outbound

This paper cites Attention Is All You Need.

Deepfake Detection of Singing Voices With Whisper Encodings Attention Is All You Need

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T22:01:29.654338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:01:29.654338Z digest=sha256:ec2245022f46d6b12a0c55a992a0a271362e07622c95d2915620edf53bff1f89

Observation 1ba5303e-4e87-4166-beb2-8fc9a90a7da8 · outbound

This paper cites Invariant representations for noisy speech recognition,.

Deepfake Detection of Singing Voices With Whisper Encodings Invariant representations for noisy speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.933211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.657540Z digest=sha256:8a31bb28b9267823f0b0b647504ed033b90878a788512dfb725f19dbc7efc744

Observation 46a6252e-bcc1-4b11-8e32-979ce0d63482 · outbound

This paper cites Learning noise-invariant representations for robust speech recognition,.

Deepfake Detection of Singing Voices With Whisper Encodings Learning noise-invariant representations for robust speech recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.922339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.660236Z digest=sha256:712ab4f2975512472e058f8f1de58ffb7975accd8b39882a290735e27658c7a7

Observation d0e43df8-43b4-4eeb-abcf-53661500b285 · outbound

This paper cites A noise-robust self-supervised pre-training model based speech representation learning for automatic speech recognition,.

Deepfake Detection of Singing Voices With Whisper Encodings A noise-robust self-supervised pre-training model based speech representation learning for automatic speech recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.911605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.663171Z digest=sha256:4919f4a5ec38ce861c821a5962cb34cc0495700ccc73a5f4593a41227ca87d30

Observation 6d92f38f-6561-4450-9781-ea6bb4005d2e · outbound

This paper cites Unsupervised learning of time–frequency patches as a noise-robust representation of speech,.

Deepfake Detection of Singing Voices With Whisper Encodings Unsupervised learning of time–frequency patches as a noise-robust representation of speech,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.900155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.666193Z digest=sha256:b276b5718d675081635ad9a3bb6908dbd83a163d4ec6a9aecae8c10180d095a1

Observation c629cc56-827a-4de2-b089-638f5822f1a2 · outbound

This paper cites Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed.

Deepfake Detection of Singing Voices With Whisper Encodings Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:01:29.669479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:01:29.669479Z digest=sha256:da402f6fc247e94c0bb7a5ba3a9eb11b4ab675be5c86c75836f329fcb100ddfe

Observation e410df5b-a605-4a62-8673-6709c5502b9a · outbound

This paper cites Computationally-efficient voice activity detection based on deep neural networks,.

Deepfake Detection of Singing Voices With Whisper Encodings Computationally-efficient voice activity detection based on deep neural networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.890043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.673284Z digest=sha256:4a4cdce9da12be758576fc4e44d56ebb53e132d855857a5b2a0f13d2923ac1ce

Observation d9913635-5c40-472f-98cf-451174f351c7 · outbound

This paper cites Pyannote. audio: neu- ral building blocks for speaker diarization,.

Deepfake Detection of Singing Voices With Whisper Encodings Pyannote. audio: neu- ral building blocks for speaker diarization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.880003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-09T22:01:29.676332Z digest=sha256:57032a3c8afb958b564cd278693171a51276c20d743831a42ade9e6891cabbb9

Pith citing papers

No inbound Pith citation observations are available.