Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:09:27.471742Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2506.03364.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:09:27.471742Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:09:24.996793Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T11:09:27.678728Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bf5d19be-3b09-4468-8649-365e3b56b752 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d3a3e3e5-0f49-4091-849a-5f3dfe8cce50 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Speech Foundation Models : We consider WavLM1 [14] and Unispeech-SAT2 [15] which are SOTA SFMs in SUPERB
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 679dcff9-5944-4529-bb9a-285cd88e275a · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models We implemented two distinct downstream for individual FMs—Fully Connected Network (FCN) and CNN
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dbe3c320-203c-41a4-83af-67f9d25ce94d · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Dataset We utilized the CtrSVDD [24], a benchmark dataset specifically designed for SVDD and the audio samples are in Chinese and Japanese
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8d79b767-cba9-4e32-9522-8ace14960afe · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Table 2 presents the evaluation scores for modeling with various combinations of SFMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f0c74d4a-3202-48ab-88a9-c0cc5e74d64b · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models MMFMs such as IB and LB, excel in capturing source-specific traits like timbre, pitch manipulation, and synthesis artifacts due to their cross- modal pretraining
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 68bce3c9-6dce-43cf-9ec2-b63feadaeb2c · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Singfake: Singing voice deepfake detection,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b5eb863e-249c-4657-868d-3723ece30843 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Ctrsvdd: A benchmark dataset and baseline analysis for controlled singing voice deepfake detection,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1048414a-db4e-4081-a730-fa67b6abafa5 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Svdd 2024: The inaugural singing voice deepfake detection chal- lenge,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4f667a30-42b5-4b98-abd0-2f903da00376 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Source tracing: Detect- ing voice spoofing,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 591a3f07-e0c7-413e-aaa0-759cb463d0ea · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models An initial investigation for detecting vocoder fingerprints of fake audio,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b6b559f-fe18-4f87-90ef-964bf6359719 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Distinguish- ing neural speech synthesis models through fingerprints in speech waveforms,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d0a2692a-b86d-4811-8d03-098f21af0a81 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Attacker attribution of audio deepfakes,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 52489c2d-efab-442d-93eb-61e994244288 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Source tracing of audio deepfake systems,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 478a36bd-6600-4e79-8918-c11b2d2d8ee7 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Attribu- tion of diffusion based deepfake speech generators,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d6129c11-13f2-48a8-a410-d86cc4314456 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Heterogeneity over homogeneity: Investigating multilingual speech pre-trained models for detecting audio deepfake,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb039946-97e3-4a0b-a850-d9025936e263 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Singing voice graph modeling for singfake detection,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d8723da1-61c5-4899-adbf-30b6a2428b1f · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Investi- gation of ensemble features of self-supervised pretrained models for automatic speech recognition,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb87ff66-0a6a-4a5c-95f1-95a181ee0a21 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Speech foundation model ensembles for the controlled singing voice deep- fake detection (ctrsvdd) challenge 2024,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9dc6690b-1a82-42a6-8356-8b71f414ec60 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ac6c5fb-6d22-4294-9b4c-19724ceed1d5 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Unispeech-sat: Universal speech repre- sentation learning with speaker aware pre-training,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb05424c-e5a7-46f8-b6be-cbeef6e447cb · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Xls-r: Self-supervised cross-lingual speech representation learning at scale,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fba6c34-fec5-41e4-a3c7-4dfdc2b7c4e4 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Robust speech recognition via large-scale weak supervision,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46591a55-2116-4698-8ede-2034bc7c413a · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Scaling speech technology to 1,000+ languages,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 117de633-59f5-4f2e-a612-26b40c75f1d7 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models X-vectors: Robust dnn embeddings for speaker recognition,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4069a85-3333-437f-b636-a2c8f78ba737 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83bd0d5-9e9d-4d74-bb6a-96484dde7067 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models MAP-Music2Vec: A Simple and Effective Baseline for Self-Supervised Music Audio Representation Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97af3196-4099-4190-88ef-d456c4532621 · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Imagebind: One embedding space to bind them all,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 425a999f-4e97-46bf-aaea-b4196c0982ac · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92b3bf40-c88b-4494-8d00-2fa2570df64e · outbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Svdd challenge 2024: A singing voice deepfake detection challenge (ctrsvdd track, training/development set),
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5d19be-3b09-4468-8649-365e3b56b752 · inbound
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.