Pith. sign in

Paper Citation Record · LEDGER

Deepfake Audio Detection Using Self-supervised Fusion Representations

As of 4 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2605.03420.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.03420 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T13:01:32.716126Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T05:19:26.613301Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T07:47:45.160739Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ea2c818-999d-45e3-bbbb-2fd72c4b1daa · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions.

Deepfake Audio Detection Using Self-supervised Fusion Representations Natural tts synthesis by conditioning wavenet on mel spectrogram predictions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.548743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:d5b798e72826816c565b9187318e2b9ae5c68df34a9843658323988bf367a539

Observation b68674b3-debb-4467-b1c0-36600ad014e6 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

Deepfake Audio Detection Using Self-supervised Fusion Representations Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.560451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:cf1b0c231cbbbffea30685fb4a7b52f8f5a198e72264fceb3487c42ca14ff139

Observation 5d9c8094-eefb-47c5-9e80-01198b8bb346 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech lan- guage models.

Deepfake Audio Detection Using Self-supervised Fusion Representations Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech lan- guage models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.585169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:af1cf60778f59def29f580c504a36dd770ef5d7b2b71838c169b196a15120f9b

Observation 59bec08d-c133-43f9-9735-7d7a5f3d37de · outbound

This paper cites Yourtts: Towards zero- shot multi-speaker tts and zero-shot voice conversion for everyone.

Deepfake Audio Detection Using Self-supervised Fusion Representations Yourtts: Towards zero- shot multi-speaker tts and zero-shot voice conversion for everyone

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.563914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:5c4baad165085d3b5eba759edd0aaf946e9f65162c3460128c02661477c50eec

Observation e53a4fcf-fe72-407d-9569-09b1965794b5 · outbound

This paper cites Starganv2-vc: A diverse, unsupervised, non-parallel framework for natural-sounding voice conversion.

Deepfake Audio Detection Using Self-supervised Fusion Representations Starganv2-vc: A diverse, unsupervised, non-parallel framework for natural-sounding voice conversion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.550229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:96c0827da8751a4ac616968e15090ba2d3dc0ef820c48c3ccd21784cfc45dbb8

Observation 43518327-1a30-4045-b1b6-0c45ff6a2772 · outbound

This paper cites Asvspoof 2019: Future horizons in spoofed and fake audio detection.

Deepfake Audio Detection Using Self-supervised Fusion Representations Asvspoof 2019: Future horizons in spoofed and fake audio detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.524557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:6e40f9a562919a49d653fd1b83f67ce4af3af86396fa613c88a64e9af1645a2d

Observation fbf3f83c-8f46-4300-8606-f763c9b2afd6 · outbound

This paper cites Spoofceleb: Speech deepfake detection and sasv in the wild.

Deepfake Audio Detection Using Self-supervised Fusion Representations Spoofceleb: Speech deepfake detection and sasv in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.532381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:bd17712e483030a05165f13eecf21cd397af66bf298689dc9ef13e42919c5729

Observation 7a6a7310-2111-4790-ac26-ca398293bd97 · outbound

This paper cites Partial fake speech attacks in the real world using deepfake audio.

Deepfake Audio Detection Using Self-supervised Fusion Representations Partial fake speech attacks in the real world using deepfake audio

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.502817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:ccb9db9672ad6f63c0f391d2bf9e9b586856728e12ee1bf231cf2e018c163f54

Observation cfccc371-3d31-4eb8-9b73-315ce904f298 · outbound

This paper cites ASVspoof 2021: Automatic Speaker Verification Spoofing and Countermeasures Challenge Evaluation Plan.

Deepfake Audio Detection Using Self-supervised Fusion Representations ASVspoof 2021: Automatic Speaker Verification Spoofing and Countermeasures Challenge Evaluation Plan

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:31.921992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:9b7a7284df8a63b37b9114ccad3ad5e2af8d2d4dc653d90062d583df1bdcf6ec

Observation aa3c082f-a97b-402a-a1aa-feb01b7ab5d7 · outbound

This paper cites Add 2022: the first audio deep synthesis detection challenge.

Deepfake Audio Detection Using Self-supervised Fusion Representations Add 2022: the first audio deep synthesis detection challenge

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.519930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:3ae3ea20c482649e8385c90eb3356d67dd149351d911e99d97d01cd4a5a40f10

Observation c4bc1d59-da12-4826-b9a8-9be29335759b · outbound

This paper cites Wavefake: A data set to facilitate audio deepfake detection.

Deepfake Audio Detection Using Self-supervised Fusion Representations Wavefake: A data set to facilitate audio deepfake detection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.560130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:8202428c87102feae5d55e796e82a97eda2e8162865b9fa9d7774d133db0c08f

Observation ac3effc8-f2d1-44b5-9c3b-3daf941455db · outbound

This paper cites Deepfake audio detection via mfcc features using machine learning.

Deepfake Audio Detection Using Self-supervised Fusion Representations Deepfake audio detection via mfcc features using machine learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.552937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:c630ad0347ed67bcee31635e3037689ac9779558321646d6b3c05304bfc708a6

Observation 20eae8f4-0122-4d03-90a2-c754f4111757 · outbound

This paper cites Hybrid transformer architectures with diverse audio features for deepfake speech classification.

Deepfake Audio Detection Using Self-supervised Fusion Representations Hybrid transformer architectures with diverse audio features for deepfake speech classification

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.538725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:64db1e0278c21922caf33aa69b79471f3fe96aed414ce42440e6846a8c615777

Observation f3da48f1-dc6c-4a8d-9255-37bc8cafc205 · outbound

This paper cites Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech.

Deepfake Audio Detection Using Self-supervised Fusion Representations Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.576381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:4043bcab93f09eb816a3bec92af5a2346c94b4e2a87b2871482a40e7547773d0

Observation f7fcc6bf-2a14-43a0-84e2-25ba48ec8ce0 · outbound

This paper cites Deep attention mechanism residual network for audio deepfake detection using multi scale cepstral coeffi- cient features.

Deepfake Audio Detection Using Self-supervised Fusion Representations Deep attention mechanism residual network for audio deepfake detection using multi scale cepstral coeffi- cient features

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.580052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:28e58fd11769f1af2ee07337ec46a6b73f7f36a3cba618befd9cdf347a05dba6

Observation 98fd61d8-92af-46d4-91f1-9fc5e835cbbf · outbound

This paper cites Mixture of experts fusion for fake audio detection using frozen wav2vec 2.0.

Deepfake Audio Detection Using Self-supervised Fusion Representations Mixture of experts fusion for fake audio detection using frozen wav2vec 2.0

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.588793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:c2f8939b7408e01eee73c2d1eb62ac47114966186a8aee45368f37696f95d7fa

Observation df7abe2e-8a29-48af-a38c-d6080be278d5 · outbound

This paper cites V oice deepfake detection using the self-supervised pre-training model hubert.

Deepfake Audio Detection Using Self-supervised Fusion Representations V oice deepfake detection using the self-supervised pre-training model hubert

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.556524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:25bb46efa0d8c11c9035e57df7ac20a0e0cb2c8a835bd9ba5d2a7ff804fc7204

Observation 18bd9f23-c802-4aec-a334-7cebf72ecca6 · outbound

This paper cites Wavlm model ensemble for audio deepfake detection.

Deepfake Audio Detection Using Self-supervised Fusion Representations Wavlm model ensemble for audio deepfake detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.568034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:ab6b3406512c3396d4b00ffeecffbb5d7bc14b25f8c2aa56f7eca760a497c3d7

Observation a577f3c7-0179-4cea-85be-1b70d59c5a95 · outbound

This paper cites Training-free deepfake voice recognition by leveraging large-scale pre-trained models.

Deepfake Audio Detection Using Self-supervised Fusion Representations Training-free deepfake voice recognition by leveraging large-scale pre-trained models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.546731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:af150b7906189581d5f78b784b843b8ffd2be90d0b6c903bc5f1e1d8b00547aa

Observation cbc48424-f100-4106-838d-671dc68cdd60 · outbound

This paper cites Xls-r: Self-supervised cross-lingual speech representation learning at scale.

Deepfake Audio Detection Using Self-supervised Fusion Representations Xls-r: Self-supervised cross-lingual speech representation learning at scale

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.540548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:2cd86124d8485ff8475bb31569163d24800ede25b4b98649fd26258266158ecf

Observation 5e6ea013-9a72-4a62-b86a-71d8d20cbf61 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers.

Deepfake Audio Detection Using Self-supervised Fusion Representations Beats: Audio pre-training with acoustic tokenizers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.553505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:223745d4b7fbd9bd7908d35a2b1bfff9ecb54d67ea709357ae40cf1902bf2592

Observation a7bb388f-d933-42ce-a3d1-027858ab98b7 · outbound

This paper cites Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks.

Deepfake Audio Detection Using Self-supervised Fusion Representations Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.572483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:dfcd0f4ce047f4432d0500c5a59bcdaace070fb58bd20d5e9d7bcda691a0b05e

Observation 7b9ac874-da98-4c95-9207-5540316c43d5 · outbound

This paper cites Esdd2: Environment-aware speech and sound deepfake detection challenge evaluation plan.

Deepfake Audio Detection Using Self-supervised Fusion Representations Esdd2: Environment-aware speech and sound deepfake detection challenge evaluation plan

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T11:39:04.523467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:01:32.716126Z digest=sha256:5e6a9a333d9308471898185f8c2992ac7b539e04eee12787f34fff5925af7ffe

Pith citing papers

Observation fab6b24a-b7fc-4ed9-aea0-e699b9a0a007 · inbound

Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge cites this paper.

Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge Deepfake Audio Detection Using Self-supervised Fusion Representations

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T07:47:45.162298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T11:44:37.193518Z digest=sha256:d3d83cda6fe1bac12f5f3e3968a3e48c299821dab818d6fcad5af6cce383fc2a

Observation 45700f7c-301d-4e92-8be6-0916ff1e4497 · inbound

Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge cites this paper.

Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge Deepfake Audio Detection Using Self-supervised Fusion Representations

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:13:45.451160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T05:19:26.613301Z digest=sha256:27b91db48829a6a53cbfdaade76a8b295807b600f53588031172c40511a4265f