Pith. sign in

Paper Citation Record · LEDGER

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.16369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16369 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:50:10.595745Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 487ed536-8ec9-4116-a1b7-b26693617478 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.442744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.442744Z digest=sha256:da3bdd39a4b470c7b21e34e865436257c1d8814429a9c09c32c4c781ac42319a

Observation bad70afb-50ed-47ed-af2a-58a02092ae6b · outbound

This paper cites Esdd2: Environment-aware speech and sound deepfake detection challenge evaluation plan,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Esdd2: Environment-aware speech and sound deepfake detection challenge evaluation plan,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.527062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.527062Z digest=sha256:4fb6d79b48cbd2808c679eb5af750c99f1c8515752fabcc68e0c3d215591d76e

Observation b3c436c7-fa36-4353-9895-f2c29d5ec2c5 · outbound

This paper cites Compspoof: A dataset and joint learning framework for component- level audio anti-spoofing countermeasures,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Compspoof: A dataset and joint learning framework for component- level audio anti-spoofing countermeasures,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.617801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.617801Z digest=sha256:9f32265961aadc3a6b7d484251f68e35b97a8106a62621f3a6134ea37be9b49d

Observation 0aff3c87-83fa-47d1-9fb0-6666bbb7709f · outbound

This paper cites Esdd2-compspoof-v2: A compos- ite spoofing dataset for speech anti-spoofing,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Esdd2-compspoof-v2: A compos- ite spoofing dataset for speech anti-spoofing,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.678153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.678153Z digest=sha256:df91903bfbfd79cc301a4162d8b470788ab3879e078104fcda5e30517ed59d09

Observation 4935244a-c2a2-4dfb-b16e-03d91b2e0c9a · outbound

This paper cites Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.772672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.772672Z digest=sha256:0c6627e133bcd6c4f4de6983665d2dca75b1d2c450b84ba94832c13f3107878c

Observation 8cbaa82c-352c-4c73-b659-4970c74494ed · outbound

This paper cites Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.861578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.861578Z digest=sha256:8560974bf07421b2178b03c1b27e15c7a4852b216883dbb3d468255b94bce06b

Observation 568f3c70-98ef-492a-a6d6-4eef0189b940 · outbound

This paper cites ASVspoof 5: crowdsourced speech data, deepfakes, and adversarial attacks at scale,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection ASVspoof 5: crowdsourced speech data, deepfakes, and adversarial attacks at scale,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.951080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.951080Z digest=sha256:3cd7fa6eeaf6e20acf26bf4837372c6021c43ac020badcf981738f6c2df59f1d

Observation 02661d5b-ae33-4204-a2c4-0b40ae623dd2 · outbound

This paper cites Towards end- to-end synthetic speech detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Towards end- to-end synthetic speech detection,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.038204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.038204Z digest=sha256:987ed214f02c0c1f019e613d180faa4a4e13c244d16f76decc4e5ad8163e2783

Observation 47682521-2c76-49fd-9d8a-98f4ddaa23af · outbound

This paper cites End-to-end anti-spoofing with rawnet2,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection End-to-end anti-spoofing with rawnet2,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.103326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.103326Z digest=sha256:3439da9b282674a3c09902ef68891232ad29bc4af89e52cdcb24508f849f208f

Observation d73820a9-44ce-439b-a063-ce4d1cf2d4af · outbound

This paper cites Does audio deepfake detection generalize?,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Does audio deepfake detection generalize?,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.193229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.193229Z digest=sha256:30a2c3f1f20f30920e0f58c7fba0f041a02c99a6cc2d3c7a408eed1e67b4f365

Observation 92d11b2f-e594-431c-8d0c-d72a58f5e7cb · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.284925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.284925Z digest=sha256:00170800543ce4862eae58a85b43723925573131774adab792bd1918733f5bfd

Observation 0f223557-e1d9-4406-a15a-2327549c8fe1 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.379834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.379834Z digest=sha256:5d1771a45d8dd147f44d5c5b7fc142241c00e1bf8f44d3dc039c1b8524025045

Observation eac88cbc-3d57-4b8f-a35a-d5e5035d938e · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.470552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.470552Z digest=sha256:a47357e7025cd61603680861de610f81c30acd9441d3a1f4b7d12fc0d8982e36

Observation 712f3198-ba83-41b6-9c1a-719c2e405483 · outbound

This paper cites Automatic speaker verification spoof- ing and deepfake detection using wav2vec 2.0 and data augmentation,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Automatic speaker verification spoof- ing and deepfake detection using wav2vec 2.0 and data augmentation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.576625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.576625Z digest=sha256:1a0514bb21b2cd9c47006f193ddfdd1374fd9eb75c74b7dde498404b5acb8345

Observation 66a6ecf3-26b0-44bb-8843-28039d2128f8 · outbound

This paper cites Xlsr-mamba: A dual-column bidirectional state space model for spoofing attack detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Xlsr-mamba: A dual-column bidirectional state space model for spoofing attack detection,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.661246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.661246Z digest=sha256:e61ac9b9e92f13499cbf6ff6c9be0afc27ef8a17dda94b34e5065c2391186d00

Observation bde7d441-5779-4fab-81a3-cd8960ef6ff0 · outbound

This paper cites Audio deepfake detection with self-supervised xls-r and sls classifier,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Audio deepfake detection with self-supervised xls-r and sls classifier,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.755238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.755238Z digest=sha256:ddde25ccf144ac8ccc235c4fe2ccfa2cdce9901a53c3ac275483bd55463233c7

Observation 4185632b-26bc-41d4-9b16-00126d402778 · outbound

This paper cites Temporal-channel modeling in multi-head self-attention for synthetic speech detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Temporal-channel modeling in multi-head self-attention for synthetic speech detection,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.822734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.822734Z digest=sha256:7946f99fa758493b569c879e8bcdde691d49f19df5a4c87cfdbf485812f45bb6

Observation 99eed178-2665-4e48-8b41-81a80db76a7a · outbound

This paper cites Scenefake: An initial dataset and benchmarks for scene fake audio detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Scenefake: An initial dataset and benchmarks for scene fake audio detection,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.909457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.909457Z digest=sha256:24098c01ca11600e5913e825d697263a16fad6292acfff1210f79ff8f4d7e7f8

Observation 169c12a5-6426-4f0b-bed0-dc6bb2016f0f · outbound

This paper cites Envfake: An initial environmental-fake audio dataset for scene-consistency detec- tion,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Envfake: An initial environmental-fake audio dataset for scene-consistency detec- tion,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:09.973741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:09.973741Z digest=sha256:c657f57d9045067c7fd5e330e0760886e05c9a08a915da9121e9ea6ed1d849f7

Observation 2c017e1e-1efe-444f-9bae-6a22cedcae5a · outbound

This paper cites Detection of deepfake environmental audio,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Detection of deepfake environmental audio,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.034929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.034929Z digest=sha256:20b54694c827bcdb46c324f9580a2d59a471d433a65c0b6b1951ac8ace0ed3a9

Observation ef1a24a4-5fc6-488d-bdea-5780b82badd3 · outbound

This paper cites Fakesound: Deepfake general audio detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Fakesound: Deepfake general audio detection,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.099628Z digest=sha256:498dbffd84f4892dc5db9a061fb0e9066f469e9bc305a5609ea454b7ebfb85e2

Observation 594f6339-d016-4a8b-9800-bc75b6c0c40a · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Audiocaps: Generating captions for audios in the wild,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.152336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.152336Z digest=sha256:4d2012ea0dde7134bd13e14f519956c8061baeecbbbe5eefe038fa5e6839f7bf

Observation e2ca0e45-ef25-49ed-9fd5-fedf4e263baf · outbound

This paper cites Envsdd: Benchmarking envi- ronmental sound deepfake detection,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Envsdd: Benchmarking envi- ronmental sound deepfake detection,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.227519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.227519Z digest=sha256:b1e3df594fb34fcfa2afe052fac0b555632ae3bf6f9d7be617754efbce57cbf5

Observation 930da6ba-77c3-4cf0-a520-00c07923907d · outbound

This paper cites Does audio deepfake detection rely on artifacts?,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Does audio deepfake detection rely on artifacts?,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.305054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.305054Z digest=sha256:d493f1aa283b4f29c6dea4744c93041fe38d01fb029655351775fcbe45dfe07c

Observation c2266bcd-6fa3-4143-bd0b-39e7c8938af8 · outbound

This paper cites Audio deepfake detection under post-processing attack,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Audio deepfake detection under post-processing attack,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.370933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.370933Z digest=sha256:8978a138ea79550a9aeea5c7b392a247aeb4bc031dc01a44410e00376904d38b

Observation 8ecec13e-a28f-420a-82d7-2f3d3d497a16 · outbound

This paper cites Do compact ssl backbones matter for audio deepfake detection? a controlled study with raptor,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Do compact ssl backbones matter for audio deepfake detection? a controlled study with raptor,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.422630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.422630Z digest=sha256:14955be48471c79b884ec80af4b3f4f8165decaa66f5fd1de752b1068f12fe2c

Observation cd0414f1-b199-455e-a042-008979f2bdc2 · outbound

This paper cites Rawboost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Rawboost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.503091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.503091Z digest=sha256:2ab6a88fc8bf705b5a9c847deb146cad90985c1cd80fa9909942bb47c982c226

Observation c79fadc9-faea-4b82-95db-f2d93668d679 · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Decoupled weight decay regulariza- tion,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:10.595745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:10.595745Z digest=sha256:f4d8b821164281f0cd63ab6f9925c21982430246773a39acae3e08e7e57d6f89

Pith citing papers

No inbound Pith citation observations are available.