Pith. sign in

Paper Citation Record · LEDGER

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection

As of 10 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.16369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16369 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:50:10.595745Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 487ed536-8ec9-4116-a1b7-b26693617478 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.442744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.442744Z digest=sha256:21030fc874ceb5484e0c78be78525451536eee37201529b2e9705db41568891d

Observation bad70afb-50ed-47ed-af2a-58a02092ae6b · outbound

This paper cites Esdd2: Environment-aware speech and sound deepfake detection challenge evaluation plan,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Esdd2: Environment-aware speech and sound deepfake detection challenge evaluation plan,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.527062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.527062Z digest=sha256:0563e8f544c5007983cab5d542bec1d4981daadecd48f331174ac87b5ded46b4

Observation b3c436c7-fa36-4353-9895-f2c29d5ec2c5 · outbound

This paper cites Compspoof: A dataset and joint learning framework for component- level audio anti-spoofing countermeasures,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Compspoof: A dataset and joint learning framework for component- level audio anti-spoofing countermeasures,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.617801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.617801Z digest=sha256:4ee15c24c6b1882b0381e25ff8f3f682bd7d8f633eba930f1ad118b00f444fec

Observation 0aff3c87-83fa-47d1-9fb0-6666bbb7709f · outbound

This paper cites Esdd2-compspoof-v2: A compos- ite spoofing dataset for speech anti-spoofing,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Esdd2-compspoof-v2: A compos- ite spoofing dataset for speech anti-spoofing,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.678153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.678153Z digest=sha256:6aa1c3fc0e6bb3b79f795b830bbe00cc8190d151d9184a693ea644826763bbb8

Observation 4935244a-c2a2-4dfb-b16e-03d91b2e0c9a · outbound

This paper cites Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.772672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.772672Z digest=sha256:586a2d7e3beb834d1e4bddb481576d3f6c2883444c921f7f35dfdd4a12be01d2

Observation 8cbaa82c-352c-4c73-b659-4970c74494ed · outbound

This paper cites Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.861578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.861578Z digest=sha256:7b05ad3c83d4b56ae00be928289631ec05e9331cd69e96393e59dcf6415d6144

Observation 568f3c70-98ef-492a-a6d6-4eef0189b940 · outbound

This paper cites ASVspoof 5: crowdsourced speech data, deepfakes, and adversarial attacks at scale,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection ASVspoof 5: crowdsourced speech data, deepfakes, and adversarial attacks at scale,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.951080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.951080Z digest=sha256:c25e5186bc65b4a4c3f2c62f7a010a7d29526ef2686fc4d3110aa77aef0d2e11

Observation 02661d5b-ae33-4204-a2c4-0b40ae623dd2 · outbound

This paper cites Towards end- to-end synthetic speech detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Towards end- to-end synthetic speech detection,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.038204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.038204Z digest=sha256:5b618565a5123cc559b150ffd3c6aa878c644f92d0b108aed0298ad0e913551f

Observation 47682521-2c76-49fd-9d8a-98f4ddaa23af · outbound

This paper cites End-to-end anti-spoofing with rawnet2,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection End-to-end anti-spoofing with rawnet2,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.103326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.103326Z digest=sha256:4ac7eede45f68291b8d787fad3dbd453665c58b09a3c8ccc3fb50684cb62bd18

Observation d73820a9-44ce-439b-a063-ce4d1cf2d4af · outbound

This paper cites Does audio deepfake detection generalize?,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Does audio deepfake detection generalize?,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.193229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.193229Z digest=sha256:47c5c0a874f6aa717a7a4764e39f4886d0c7867e27ed12ce542a79b1af054a45

Observation 92d11b2f-e594-431c-8d0c-d72a58f5e7cb · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.284925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.284925Z digest=sha256:9ea03ff48fbb7e8e7ddcc333a14e2f6ca89520e1953b98cce588de9faf615fbc

Observation 0f223557-e1d9-4406-a15a-2327549c8fe1 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.379834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.379834Z digest=sha256:4e0adb6e0725696b2826921a6f35a3f05299d1ed2685f6c9b71b11a0a8acf4ad

Observation eac88cbc-3d57-4b8f-a35a-d5e5035d938e · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.470552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.470552Z digest=sha256:0ce99a8b8ec93ecd0bf32bf3d42c78a2ed6562dd4c1056f1180bc83cd28e75fc

Observation 712f3198-ba83-41b6-9c1a-719c2e405483 · outbound

This paper cites Automatic speaker verification spoof- ing and deepfake detection using wav2vec 2.0 and data augmentation,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Automatic speaker verification spoof- ing and deepfake detection using wav2vec 2.0 and data augmentation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.576625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.576625Z digest=sha256:5def6392563073e8e232d6ec43205a68a3ddef88a503ad003c496fa63aeb34ce

Observation 66a6ecf3-26b0-44bb-8843-28039d2128f8 · outbound

This paper cites Xlsr-mamba: A dual-column bidirectional state space model for spoofing attack detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Xlsr-mamba: A dual-column bidirectional state space model for spoofing attack detection,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.661246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.661246Z digest=sha256:f50e4ba01c8c9ecee7a7716409480629e64e9c68fd1a06f350dac8bb3008a0bd

Observation bde7d441-5779-4fab-81a3-cd8960ef6ff0 · outbound

This paper cites Audio deepfake detection with self-supervised xls-r and sls classifier,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Audio deepfake detection with self-supervised xls-r and sls classifier,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.755238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.755238Z digest=sha256:363b34190e1fa15e8bed21437fb0138bbbde98a37fab34f1c6622c9560ba7f0f

Observation 4185632b-26bc-41d4-9b16-00126d402778 · outbound

This paper cites Temporal-channel modeling in multi-head self-attention for synthetic speech detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Temporal-channel modeling in multi-head self-attention for synthetic speech detection,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.822734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.822734Z digest=sha256:eb97813963937bb61087bd1b067a8af60fe6ebeafdbcec7a6420b7ef422d94a9

Observation 99eed178-2665-4e48-8b41-81a80db76a7a · outbound

This paper cites Scenefake: An initial dataset and benchmarks for scene fake audio detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Scenefake: An initial dataset and benchmarks for scene fake audio detection,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.909457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.909457Z digest=sha256:b92b598534466b0bbf7c0dcd5015fa06b34c7e01fb39aae56824cd0209e576cb

Observation 169c12a5-6426-4f0b-bed0-dc6bb2016f0f · outbound

This paper cites Envfake: An initial environmental-fake audio dataset for scene-consistency detec- tion,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Envfake: An initial environmental-fake audio dataset for scene-consistency detec- tion,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.973741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.973741Z digest=sha256:25c24a7faf48f177a62e8dbc3e67c36c0325a593ad825fb035baec5ef19efde4

Observation 2c017e1e-1efe-444f-9bae-6a22cedcae5a · outbound

This paper cites Detection of deepfake environmental audio,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Detection of deepfake environmental audio,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.034929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.034929Z digest=sha256:49f4ea0be1e372ac829fa060bcb2dcc314fbdde19c9b2715145ecead9c3fa230

Observation ef1a24a4-5fc6-488d-bdea-5780b82badd3 · outbound

This paper cites Fakesound: Deepfake general audio detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Fakesound: Deepfake general audio detection,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.099628Z digest=sha256:37ec21f422f9f090ce5502d3908d6e592f687702de0715eb59e466727765821a

Observation 594f6339-d016-4a8b-9800-bc75b6c0c40a · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Audiocaps: Generating captions for audios in the wild,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.152336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.152336Z digest=sha256:7a72945f7c47567d25dfe2f57db14c2f9726f99faebcef928e289d51a1fedbb5

Observation e2ca0e45-ef25-49ed-9fd5-fedf4e263baf · outbound

This paper cites Envsdd: Benchmarking envi- ronmental sound deepfake detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Envsdd: Benchmarking envi- ronmental sound deepfake detection,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.227519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.227519Z digest=sha256:1252b741df365dee1d01aa09dcd4600ae1effc09114268d5510a44ba68fdc441

Observation 930da6ba-77c3-4cf0-a520-00c07923907d · outbound

This paper cites Does audio deepfake detection rely on artifacts?,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Does audio deepfake detection rely on artifacts?,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.305054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.305054Z digest=sha256:fb4403ecc2afa8cf7fadfcb6892506b4b919dead634b0a8d50ba132c7e7068a3

Observation c2266bcd-6fa3-4143-bd0b-39e7c8938af8 · outbound

This paper cites Audio deepfake detection under post-processing attack,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Audio deepfake detection under post-processing attack,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.370933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.370933Z digest=sha256:127d5c448f7d74f1dad8e8fa7352e670418998aab3280fadd93151c68671cb76

Observation 8ecec13e-a28f-420a-82d7-2f3d3d497a16 · outbound

This paper cites Do compact ssl backbones matter for audio deepfake detection? a controlled study with raptor,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Do compact ssl backbones matter for audio deepfake detection? a controlled study with raptor,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.422630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.422630Z digest=sha256:91466463e721787fe49f4053588640476752e5c8c1ad593dee5ff4faeeba47aa

Observation cd0414f1-b199-455e-a042-008979f2bdc2 · outbound

This paper cites Rawboost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Rawboost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.503091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.503091Z digest=sha256:86c60185130b0a5acc086ce15a38ad7856a4a7a0eeee0f2818a97160d7355c7d

Observation c79fadc9-faea-4b82-95db-f2d93668d679 · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Decoupled weight decay regulariza- tion,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.595745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.595745Z digest=sha256:0dbcf680e62a35d26ec30c6655ce654dcc4f14e64bc81a06ade586b43d64c721

Pith citing papers

No inbound Pith citation observations are available.