Pith. sign in

REVIEW 2 major objections 2 minor 2 cited by

The RADAR Challenge 2026 benchmark shows that audio deepfake detectors still produce high equal error rates when tested on multilingual utterances after compression, resampling, noise, and reverberation.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 22:51 UTC pith:FI2SAV6A

load-bearing objection A standard challenge announcement paper that sets up a multilingual audio deepfake benchmark with media transformations and reports participation numbers, but stays descriptive without new methods or detailed results. the 2 major comments →

arxiv 2605.09568 v3 pith:FI2SAV6A submitted 2026-05-10 eess.AS

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations

classification eess.AS
keywords audio deepfake detectionmedia transformationsmultilingual evaluationequal error rategrand challengerobust recognitioncompression and noise
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper presents the task definition, dataset construction, and evaluation outcomes for a grand challenge on audio deepfake detection. The challenge supplies labeled English development data and then evaluates submissions on more than 100,000 utterances spanning six languages under realistic media transformations. Twenty-two teams participated in the final phase; their collective results demonstrate that current detectors do not yet achieve reliable real-versus-fake classification once the audio has passed through typical distribution pipelines.

Core claim

The central discovery is that binary real/fake classifiers evaluated by equal error rate continue to exhibit elevated error rates on the multilingual, media-transformed test set, indicating that robustness under compression, resampling, additive noise, and reverberation across languages remains an open problem.

What carries the argument

Equal error rate (EER) computed on the multilingual media-transformed evaluation utterances, which serves as the sole performance metric for ranking submissions.

Load-bearing premise

The particular choices of media transformations and language distribution in the test set accurately mirror the conditions that real-world audio deepfakes will encounter.

What would settle it

A follow-up experiment in which the top-performing systems from the challenge are re-evaluated on a fresh collection of real and synthetic audio that has passed through the same transformation pipeline but was collected independently of the challenge data.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Any production detector must maintain low EER after the audio has undergone lossy compression and resampling.
  • Performance must generalize from English to at least five additional languages without retraining on target-language fakes.
  • Noise and reverberation must be treated as first-class distortions rather than optional augmentations.
  • Future systems will need to report EER on held-out multilingual transformed data rather than clean English test sets alone.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Deployment pipelines that ingest user-uploaded audio will require additional preprocessing stages or model updates to close the observed performance gap.
  • The challenge protocol could be extended by adding a third phase that measures latency and memory use on edge devices under the same transformations.
  • Cross-lingual transfer methods developed for this task may also improve robustness in related audio forensics problems such as speaker verification under distortion.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript announces the RADAR Challenge 2026, an APSIPA Grand Challenge on robust audio deepfake recognition under media transformations (compression, resampling, noise, reverberation). It describes a two-phase structure—an English-labeled development phase and a multilingual evaluation phase with >100,000 utterances across English, Singapore English, Mandarin Chinese, Taiwanese Mandarin, Japanese, and Vietnamese—along with the EER-based binary classification protocol, participation (33 development submissions, 22 evaluation submissions), and the observation that the reported results indicate remaining challenges in multilingual and transformed conditions.

Significance. If the dataset construction and protocol are reproducible, the challenge supplies a large-scale, multilingual benchmark that can serve as a community testbed for systems intended to operate in realistic media pipelines, potentially accelerating progress on robustness beyond clean English conditions.

major comments (2)
  1. [Results / overall results paragraph] The central claim that the reported results 'highlight the remaining challenges' is load-bearing for the paper's contribution, yet the manuscript supplies no quantitative EER values, baseline comparisons, or per-language/per-transformation breakdowns to support it; without these numbers the assertion remains unsupported by evidence.
  2. [Data set construction] Dataset construction section: no information is given on how the media transformations were implemented (specific codecs, SNR ranges, reverberation parameters) or on any validation that the transformed utterances preserve the original real/fake labels, which directly affects whether the EER metric can be interpreted as measuring robustness.
minor comments (2)
  1. [Abstract] The abstract states the evaluation contains 'more than 100,000 utterances' but does not break down the counts by language or by real/fake class; adding these counts would improve clarity.
  2. [Participation summary] The paper mentions '33 teams submitted to the development phase and 22 teams submitted to the final evaluation phase' but does not indicate whether the same teams participated in both phases or whether any overlap analysis was performed.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We appreciate the referee's constructive comments on our manuscript describing the RADAR Challenge 2026. We address each major comment below and will revise the paper to strengthen the presentation of results and dataset details.

read point-by-point responses
  1. Referee: [Results / overall results paragraph] The central claim that the reported results 'highlight the remaining challenges' is load-bearing for the paper's contribution, yet the manuscript supplies no quantitative EER values, baseline comparisons, or per-language/per-transformation breakdowns to support it; without these numbers the assertion remains unsupported by evidence.

    Authors: We agree that the manuscript would benefit from explicit quantitative support for this claim. In the revised version, we will add the EER results from the top-performing teams in both phases, include a baseline system for comparison, and provide breakdowns by language and transformation type where space permits. This will allow readers to better assess the remaining challenges. revision: yes

  2. Referee: [Data set construction] Dataset construction section: no information is given on how the media transformations were implemented (specific codecs, SNR ranges, reverberation parameters) or on any validation that the transformed utterances preserve the original real/fake labels, which directly affects whether the EER metric can be interpreted as measuring robustness.

    Authors: We acknowledge that additional details are needed. The revised manuscript will provide the specific implementation parameters for the media transformations and describe the procedures used to validate that the transformations do not alter the real/fake labels. revision: yes

Circularity Check

0 steps flagged

No significant circularity; purely descriptive challenge paper

full rationale

The paper is a standard challenge announcement describing task setup, multilingual dataset construction, evaluation protocol using EER, and participant outcomes (33 dev, 22 eval submissions). It advances no derivations, equations, predictions, models, or theorems. The central claim is observational: reported EERs indicate remaining difficulties under the defined conditions. No self-citations are load-bearing for any derivation, and no steps reduce by construction to inputs or prior author work. This is self-contained against external benchmarks as a simulation tool.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

No mathematical derivations, fitted parameters, or postulated entities are involved; the paper is a challenge description.

pith-pipeline@v0.9.1-grok · 5698 in / 1058 out tokens · 20488 ms · 2026-06-30T22:51:21.032645+00:00 · methodology

0 comments
read the original abstract

RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression, resampling, noise, and reverberation. It consists of two phases: an English development phase with labeled data for analysis and paper writing, and a multilingual evaluation phase containing more than 100,000 utterances in English, Singapore English, Mandarin Chinese, Taiwanese Mandarin, Japanese, and Vietnamese. Systems are evaluated using equal error rate (EER) for binary real/fake classification. This paper describes the challenge task, the construction of the data set, the evaluation protocol, and the overall results. During the challenge, 33 teams submitted to the development phase and 22 teams submitted to the final evaluation phase. The reported results highlight the remaining challenges of robust audio deepfake detection under multilingual and media-transformed conditions.

Figures

Figures reproduced from arXiv: 2605.09568 by Hieu-Thi Luong, Ivan Kukanov, Kong Aik Lee, Xuechen Liu, Zheng Xin Chai.

Figure 1
Figure 1. Figure 1: The EER results in Phase 1 and Phase 2 of the top 26 teams. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Large Audio Language Models for Spoofing-Aware Speaker Verification

    cs.SD 2026-07 conditional novelty 6.0

    Adapted LALMs can reach competitive spoofing-aware speaker verification (89.3% accuracy, 0.19 min a-DCF on an ASVspoof5 subset), though zero-shot performance is near chance.

  2. Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English

    eess.AS 2026-07 conditional novelty 4.0

    Fine-tuning Chatterbox and CosyVoice 3 on 50 Singlish speakers measurably raises accent similarity, and the gain persists on held-out speakers.

Reference graph

Works this paper leans on

38 extracted references · 38 canonical work pages · cited by 2 Pith papers · 6 internal anchors

  1. [1]

    ASVspoof 2019: A large-scale pub- lic database of synthesized, converted and replayed speech,

    X. Wang et al., “ASVspoof 2019: A large-scale pub- lic database of synthesized, converted and replayed speech,”Computer Speech & Language, vol. 64, p. 101 114, Nov. 2020,ISSN: 08852308

  2. [2]

    Insights into deep non-linear filters for improved multi-channel speech enhancement,

    X. Liu et al., “ASVspoof 2021: Towards spoofed and deepfake speech detection in the wild,”IEEE/ACM Transactions on Audio, Speech, and Language Process- ing, vol. 31, pp. 2507–2522, 2023.DOI: 10.1109/TASLP. 2023.3285283

  3. [3]

    Lipiecki, K

    X. Wang et al., “ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and ad- versarial attack detection using crowdsourced speech,” Computer Speech & Language, vol. 95, p. 101 825, 2026,ISSN: 0885-2308.DOI: https://doi.org/10.1016/j. csl.2025.101825

  4. [4]

    Safe: Synthetic audio forensics evalua- tion challenge,

    T. Kirill et al., “Safe: Synthetic audio forensics evalua- tion challenge,” inProc. ACM IH&MMSEC Workshop, 2025, pp. 174–180

  5. [5]

    Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,

    J. Yi et al., “ADD 2022: The first audio deep synthesis detection challenge,” inProc. ICASSP, 2022, pp. 9216– 9220.DOI: 10.1109/ICASSP43922.2022.9746939

  6. [6]

    ADD 2023: The Second Audio Deepfake Detection Challenge,

    J. Yi et al., “ADD 2023: The Second Audio Deepfake Detection Challenge,” inProc. IJCAI DADA Workshop, May 2023

  7. [7]

    Perturbed public voices (p 2v): A dataset for robust audio deepfake detection,

    C. Gao, M. Postiglione, I. Gortner, S. Kraus, and V . Sub- rahmanian, “Perturbed public voices (p 2v): A dataset for robust audio deepfake detection,”arXiv preprint arXiv:2508.10949, 2025

  8. [8]

    Room impulse responses help attackers to evade deep fake detection,

    H.-T. Luong, D.-T. Truong, K. A. Lee, and E. S. Chng, “Room impulse responses help attackers to evade deep fake detection,” inProc. SLT 2024, IEEE, 2024, pp. 623–629

  9. [9]

    DeePen: Penetration Testing for Audio Deepfake Detection

    N. M ¨uller et al., “Deepen: Penetration testing for audio deepfake detection,”arXiv preprint arXiv:2502.20427, 2025

  10. [10]

    Investigating the impact of speech enhancement on audio deep- fake detection in noisy environments,

    S. Kshirsagar, A. R. Avila, et al., “Investigating the impact of speech enhancement on audio deep- fake detection in noisy environments,”arXiv preprint arXiv:2603.14767, 2026

  11. [11]

    Mlaad: The multi-language audio anti-spoofing dataset,

    N. M. M ¨uller et al., “Mlaad: The multi-language audio anti-spoofing dataset,” inProc. IJCNN 2024, IEEE, 2024, pp. 1–7

  12. [12]

    Sea-spoof: Bridging the gap in multilingual audio deepfake detection for south-east asian,

    J. Wu, N. Hou, Z. Pan, Q. Zhang, S. H. Bhupendra, and S. Mondal, “Sea-spoof: Bridging the gap in multilingual audio deepfake detection for south-east asian,”arXiv preprint arXiv:2509.19865, 2025

  13. [13]

    Llamapartialspoof: An llm-driven fake speech dataset simulating disinformation generation,

    H.-T. Luong, H. Li, L. Zhang, K. A. Lee, and E. S. Chng, “Llamapartialspoof: An llm-driven fake speech dataset simulating disinformation generation,” inProc. ICASSP, 2025.DOI: 10 . 1109 / ICASSP49660 . 2025 . 10888070

  14. [14]

    LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,

    H. Zen et al., “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,” inInterspeech 2019, 2019, pp. 2638–2642

  15. [15]

    JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech,

    J. Lim, J. Ye, S. Chun, S. Kim, and J. Cho, “JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech,” inInterspeech 2022, 2022, pp. 2338–2342

  16. [16]

    YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,

    E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G¨olge, and M. A. Ponti, “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,” inICML, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, Eds., ser. Proceed- ings of Machine Learning Research, vol. 162, PMLR, 17–23 Jul 2022, pp. 2709–2720

  17. [17]

    XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,

    E. Casanova et al., “XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,” inInterspeech 2024, 2024, pp. 4978–4982

  18. [18]

    CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

    Z. Du et al., “Cosyvoice: A scalable multilingual zero- shot text-to-speech synthesizer based on supervised se- mantic tokens,”arXiv preprint arXiv:2407.05407, 2024

  19. [19]

    Common voice: A massively- multilingual speech corpus,

    R. Ardila et al., “Common voice: A massively- multilingual speech corpus,” inProceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), 2020, pp. 4211–4215

  20. [20]

    arXiv preprint arXiv:2111.09344 , year=

    D. Galvez et al., “The people’s speech: A large-scale di- verse english speech recognition dataset for commercial usage,”arXiv preprint arXiv:2111.09344, 2021

  21. [21]

    Building the Singapore English National Speech Corpus,

    J. X. Koh et al., “Building the Singapore English National Speech Corpus,” inInterspeech 2019, pp. 321– 325

  22. [22]

    imagicdatatech.com/index.php/home/dataopensource/ data info/id/101, Accessed: 2019-05, 2019

    Magic Data Technology Co., Ltd.,Openslr68: Magic- data mandarin chinese read speech corpus, http://www. imagicdatatech.com/index.php/home/dataopensource/ data info/id/101, Accessed: 2019-05, 2019

  23. [23]

    Formosa speech recognition challenge 2020 and taiwanese across taiwan corpus,

    Y .-F. Liao et al., “Formosa speech recognition challenge 2020 and taiwanese across taiwan corpus,” inProc. O- COCOSDA 2020, IEEE, 2020, pp. 65–70

  24. [24]

    Cpjd corpus: Crowd- sourced parallel speech corpus of japanese dialects,

    S. Takamichi and H. Saruwatari, “Cpjd corpus: Crowd- sourced parallel speech corpus of japanese dialects,” in Proc. LREC 2018, 2018

  25. [25]

    D. C. Tran,FPT Open Speech Dataset (FOSD) - Viet- namese, version V4, Mendeley Data, 2020.DOI: 10 . 17632/k9sxg2twv4.4

  26. [26]

    CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

    Z. Du et al., “CosyV oice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training,”arXiv preprint arXiv:2505.17589, 2025

  27. [27]

    Qwen3-TTS Technical Report

    H. Hu et al., “Qwen3-tts technical report,”arXiv preprint arXiv:2601.15621, 2026

  28. [28]

    Fish audio s2 technical report,

    S. Liao et al., “Fish audio s2 technical report,”arXiv preprint arXiv:2603.08823, 2026

  29. [29]

    Statistics of natural reverberation enable perceptual separation of sound and space,

    J. Traer and J. H. McDermott, “Statistics of natural reverberation enable perceptual separation of sound and space,”PNAS, vol. 113, no. 48, E7856–E7865, 2016

  30. [30]

    MUSAN: A Music, Speech, and Noise Corpus

    D. Snyder, G. Chen, and D. Povey, “Musan: A music, speech, and noise corpus,”arXiv preprint arXiv:1510.08484, 2015

  31. [31]

    Fma: A dataset for music analysis,

    M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson, “Fma: A dataset for music analysis,” in18th International Society for Music Information Retrieval Conference, 2017

  32. [32]

    A binaural room im- pulse response database for the evaluation of dereverber- ation algorithms,

    M. Jeub, M. Schafer, and P. Vary, “A binaural room im- pulse response database for the evaluation of dereverber- ation algorithms,” in2009 16th international conference on digital signal processing, IEEE, 2009, pp. 1–5

  33. [33]

    Image method for ef- ficiently simulating small-room acoustics,

    J. B. Allen and D. A. Berkley, “Image method for ef- ficiently simulating small-room acoustics,”The Journal of the Acoustical Society of America, vol. 65, no. 4, pp. 943–950, 1979

  34. [34]

    Hierarchical and multimodal learning for hetero- geneous sound classification,

    P. Anastasopoulou, F. A. Dal R ´ı, X. Serra, and F. Font, “Hierarchical and multimodal learning for hetero- geneous sound classification,” inProc. DCASE 2025, 2025

  35. [35]

    Automatic Speaker Verification Spoofing and Deepfake Detection Using Wav2vec 2.0 and Data Augmentation,

    H. Tak, M. Todisco, X. Wang, J.-w. Jung, J. Yamagishi, and N. Evans, “Automatic Speaker Verification Spoofing and Deepfake Detection Using Wav2vec 2.0 and Data Augmentation,” inProc. Odyssey 2022, 2022, pp. 112– 119

  36. [36]

    Robust localization of partially fake speech: Metrics and out-of-domain evaluation,

    H.-T. Luong, I. Rimon, H. Permuter, K. A. Lee, and E. S. Chng, “Robust localization of partially fake speech: Metrics and out-of-domain evaluation,” inProc. APSIPA ASC 2025, IEEE, 2025, pp. 2205–2210

  37. [37]

    Li, X., Chen, P.-Y ., and Wei, W

    X. Li, P.-Y . Chen, and W. Wei, “Measuring the ro- bustness of audio deepfake detectors,”arXiv preprint arXiv:2503.17577, 2025

  38. [38]

    Replay attacks against audio deepfake detection,

    N. M ¨uller et al., “Replay attacks against audio deepfake detection,”Interspeech 2025, 2025