Pith. sign in

Paper Citation Record · LEDGER

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

As of 13 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2506.03364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03364 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:09:27.471742Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:09:24.996793Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:09:27.678728Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved12
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf5d19be-3b09-4468-8649-365e3b56b752 · outbound

This paper cites Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:09:27.826462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:24.996793Z digest=sha256:43a51697f3adbe11a260f7e1f5808b50ae0b40fa5320d1218fa3d623dcf5909b

Observation d3a3e3e5-0f49-4091-849a-5f3dfe8cce50 · outbound

This paper cites Speech Foundation Models : We consider WavLM1 [14] and Unispeech-SAT2 [15] which are SOTA SFMs in SUPERB.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Speech Foundation Models : We consider WavLM1 [14] and Unispeech-SAT2 [15] which are SOTA SFMs in SUPERB

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:30.324853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.045807Z digest=sha256:016418d579fc485cb8de28a36ab63a6d01aa077935e46e2c385d2ca3b58d8393

Observation 679dcff9-5944-4529-bb9a-285cd88e275a · outbound

This paper cites We implemented two distinct downstream for individual FMs—Fully Connected Network (FCN) and CNN.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models We implemented two distinct downstream for individual FMs—Fully Connected Network (FCN) and CNN

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:30.233288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.105817Z digest=sha256:337b1e40d7f135c14e904394112437b6c1a71c59f47e5525cf267ca248d27191

Observation dbe3c320-203c-41a4-83af-67f9d25ce94d · outbound

This paper cites Dataset We utilized the CtrSVDD [24], a benchmark dataset specifically designed for SVDD and the audio samples are in Chinese and Japanese.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Dataset We utilized the CtrSVDD [24], a benchmark dataset specifically designed for SVDD and the audio samples are in Chinese and Japanese

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:09:30.146119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.190869Z digest=sha256:5e64e08808744c665b5cb8a3b3eb0f12d03ffab6db4865023ccea2dfa0ebf2c1

Observation 8d79b767-cba9-4e32-9522-8ace14960afe · outbound

This paper cites Table 2 presents the evaluation scores for modeling with various combinations of SFMs.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Table 2 presents the evaluation scores for modeling with various combinations of SFMs

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:09:30.072457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.265132Z digest=sha256:97f76f694d3fa315cfdb7d102ecd34046da91ad2ad1091344d7ece81afaf9db5

Observation f0c74d4a-3202-48ab-88a9-c0cc5e74d64b · outbound

This paper cites MMFMs such as IB and LB, excel in capturing source-specific traits like timbre, pitch manipulation, and synthesis artifacts due to their cross- modal pretraining.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models MMFMs such as IB and LB, excel in capturing source-specific traits like timbre, pitch manipulation, and synthesis artifacts due to their cross- modal pretraining

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:30.002183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.326593Z digest=sha256:352a8b083f76e18a3ed185c1347d2ead74ad33bdab073dd2e36608002c080053

Observation 68bce3c9-6dce-43cf-9ec2-b63feadaeb2c · outbound

This paper cites Singfake: Singing voice deepfake detection,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Singfake: Singing voice deepfake detection,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:29.927959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.384372Z digest=sha256:c23a343441262c253cc3fa28d3e4ccdd5f43123d3544c4c59cc370052d36374d

Observation b5eb863e-249c-4657-868d-3723ece30843 · outbound

This paper cites Ctrsvdd: A benchmark dataset and baseline analysis for controlled singing voice deepfake detection,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Ctrsvdd: A benchmark dataset and baseline analysis for controlled singing voice deepfake detection,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:29.860741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.442291Z digest=sha256:d07b5ea70aef88b688e5e636b26dc115a020fb475b43f8ede149cfe1d2a304b5

Observation 1048414a-db4e-4081-a730-fa67b6abafa5 · outbound

This paper cites Svdd 2024: The inaugural singing voice deepfake detection chal- lenge,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Svdd 2024: The inaugural singing voice deepfake detection chal- lenge,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:29.805212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.500181Z digest=sha256:13a284bc5aee895de9d9f76fc9845f44cab1b86e887c56b029e64aae327b92fa

Observation 4f667a30-42b5-4b98-abd0-2f903da00376 · outbound

This paper cites Source tracing: Detect- ing voice spoofing,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Source tracing: Detect- ing voice spoofing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:29.757860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.562248Z digest=sha256:035bbc5c6ddfd5525ca92cb8bbdceacb0242ef54a5e92cad9e0ac8e503121f60

Observation 591a3f07-e0c7-413e-aaa0-759cb463d0ea · outbound

This paper cites An initial investigation for detecting vocoder fingerprints of fake audio,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models An initial investigation for detecting vocoder fingerprints of fake audio,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:29.629903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.649992Z digest=sha256:62a88d9fef386ee618054a9d25d56b167abb1c1e93567bca92355b3361a6e48c

Observation 5b6b559f-fe18-4f87-90ef-964bf6359719 · outbound

This paper cites Distinguish- ing neural speech synthesis models through fingerprints in speech waveforms,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Distinguish- ing neural speech synthesis models through fingerprints in speech waveforms,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:29.399707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.727261Z digest=sha256:493f3ebef3a3de458dbde039c3ad0a53e75c98b7a036feb36439acd0d17ed9a2

Observation d0a2692a-b86d-4811-8d03-098f21af0a81 · outbound

This paper cites Attacker attribution of audio deepfakes,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Attacker attribution of audio deepfakes,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:29.126713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.804819Z digest=sha256:9dbe686225ef3a0ff1fefb40bbd3f43f4fa82a1daa1e20e4070c573a04a9786d

Observation 52489c2d-efab-442d-93eb-61e994244288 · outbound

This paper cites Source tracing of audio deepfake systems,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Source tracing of audio deepfake systems,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:28.900095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.877776Z digest=sha256:1fe679f696fd53faccfb4e1e3d9fef78b6ac3fbde9680039dae0734729c03ad1

Observation 478a36bd-6600-4e79-8918-c11b2d2d8ee7 · outbound

This paper cites Attribu- tion of diffusion based deepfake speech generators,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Attribu- tion of diffusion based deepfake speech generators,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:28.659768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:25.951637Z digest=sha256:7b7d8867639b8297319815075c541ac3b5a38db4450343bae645ef9d641698b0

Observation d6129c11-13f2-48a8-a410-d86cc4314456 · outbound

This paper cites Heterogeneity over homogeneity: Investigating multilingual speech pre-trained models for detecting audio deepfake,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Heterogeneity over homogeneity: Investigating multilingual speech pre-trained models for detecting audio deepfake,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.013846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.013846Z digest=sha256:0ec868d6759fd85508d9c75ee10f5169d180753dfe53cabeaa410a4fd707dfb9

Observation fb039946-97e3-4a0b-a850-d9025936e263 · outbound

This paper cites Singing voice graph modeling for singfake detection,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Singing voice graph modeling for singfake detection,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:28.417239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:26.091374Z digest=sha256:7a7c0241289df7599f90ec1e2ba86f429632446bee14e345fc2919e9012bf610

Observation d8723da1-61c5-4899-adbf-30b6a2428b1f · outbound

This paper cites Investi- gation of ensemble features of self-supervised pretrained models for automatic speech recognition,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Investi- gation of ensemble features of self-supervised pretrained models for automatic speech recognition,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.157753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.157753Z digest=sha256:5057da4321ac3ce5339798e641f6d5d24d55ec045f1f552fca9819571b33e67f

Observation eb87ff66-0a6a-4a5c-95f1-95a181ee0a21 · outbound

This paper cites Speech foundation model ensembles for the controlled singing voice deep- fake detection (ctrsvdd) challenge 2024,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Speech foundation model ensembles for the controlled singing voice deep- fake detection (ctrsvdd) challenge 2024,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:28.180367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:26.215930Z digest=sha256:28d320963420f53dbc0d3a36c9bc2e26b6f26ec50b04bc8912c7a6e9d20da55c

Observation 9dc6690b-1a82-42a6-8356-8b71f414ec60 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.303298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.303298Z digest=sha256:b0d654b46c7911bd16b8b578c812ac0d6d2c9908e2df947fed529191083b1484

Observation 3ac6c5fb-6d22-4294-9b4c-19724ceed1d5 · outbound

This paper cites Unispeech-sat: Universal speech repre- sentation learning with speaker aware pre-training,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Unispeech-sat: Universal speech repre- sentation learning with speaker aware pre-training,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.366094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.366094Z digest=sha256:7cf4596c6cfb83f1a01b305e8b1e090a361a59ed84e2b6a3cb00199a24e1e1f6

Observation bb05424c-e5a7-46f8-b6be-cbeef6e447cb · outbound

This paper cites Xls-r: Self-supervised cross-lingual speech representation learning at scale,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Xls-r: Self-supervised cross-lingual speech representation learning at scale,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.435761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.435761Z digest=sha256:9e2334d890d0669ce401e415889694b5663ab40f9b37ca240483cb96eb94c0bc

Observation 7fba6c34-fec5-41e4-a3c7-4dfdc2b7c4e4 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Robust speech recognition via large-scale weak supervision,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.502914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.502914Z digest=sha256:288ddfcbef95c3a4ba973aa358d37758c5e87f0ca4306797fcd850e152f62bf6

Observation 46591a55-2116-4698-8ede-2034bc7c413a · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Scaling speech technology to 1,000+ languages,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.571595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.571595Z digest=sha256:506a452fa2ca33d3896064f630d16303ece994ac34789d516a2e4e6bd31358c1

Observation 117de633-59f5-4f2e-a612-26b40c75f1d7 · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models X-vectors: Robust dnn embeddings for speaker recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.716713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.716713Z digest=sha256:e5e3ac8f3cf7565d351c91e42c866e2d8cc6192811502b098ff34588ca0cd567

Observation e4069a85-3333-437f-b636-a2c8f78ba737 · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.894022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.894022Z digest=sha256:10813d27e081a5b41f4e23a0e88a55c9a083005d21823a3cb607316bb544e3f2

Observation d83bd0d5-9e9d-4d74-bb6a-96484dde7067 · outbound

This paper cites MAP-Music2Vec: A Simple and Effective Baseline for Self-Supervised Music Audio Representation Learning.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models MAP-Music2Vec: A Simple and Effective Baseline for Self-Supervised Music Audio Representation Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.953234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.953234Z digest=sha256:d67bc1468b3bfdf1bb9ec8f5e37c02c1c9d7c9c5e49917fbc16d01b8c7b57ab7

Observation 97af3196-4099-4190-88ef-d456c4532621 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Imagebind: One embedding space to bind them all,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:27.157036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:27.157036Z digest=sha256:a973277b952e9af4a397580bd89ffe4c978911da911e0869959b99df06ae7080

Observation 425a999f-4e97-46bf-aaea-b4196c0982ac · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:27.322111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:27.322111Z digest=sha256:b5481cde7786cf9c75e83b251c1f353eb931706694099d2b20df8b65f74896b3

Observation 92b3bf40-c88b-4494-8d00-2fa2570df64e · outbound

This paper cites Svdd challenge 2024: A singing voice deepfake detection challenge (ctrsvdd track, training/development set),.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Svdd challenge 2024: A singing voice deepfake detection challenge (ctrsvdd track, training/development set),

Reference 30

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:09:27.471742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:27.471742Z digest=sha256:435988089c03136186f0744f6e5c2b2e35cd1356287b35a79ad0769873dad51f

Pith citing papers

Observation bf5d19be-3b09-4468-8649-365e3b56b752 · inbound

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models cites this paper.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:09:27.826462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T11:09:24.996793Z digest=sha256:43a51697f3adbe11a260f7e1f5808b50ae0b40fa5320d1218fa3d623dcf5909b