{"id":"8dfbbeb1-5a67-4710-8197-0cf52fbfbf6f","arxiv_id":"2411.15457","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"HAV-DF is a new Hindi audio-video deepfake dataset of 508 videos, and the paper reports that existing detectors score lower on it than on English-language datasets.","lead":"Researchers built a new Hindi-language deepfake dataset, HAV-DF, with 508 videos combining face swap, lip-sync, and voice cloning manipulations. It is intended to help train and test deepfake detectors for Hindi content, where few resources exist.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-dataset accuracy comparison in §5.2.2 is confounded by pipeline, train-set size, and evaluation protocol; the claim that lower HAV-DF accuracy is due to Hindi language is not established.","rationale":"The reader identified the same load-bearing concern: the comparison in §5.2.2 is confounded, so the paper's central claim about Hindi-language-driven detection difficulty is not supported by the evidence as presented. This is the most load-bearing issue because it is the only quantitative evidence for the claimed 'deeper challenge' of HAV-DF. The manuscript itself supplies alternative explanations in §6.1: small dataset size, single-frontal-subject constraint, and limited manipulation types. Secondary issues reinforce the conditional verdict: dataset cardinality is inconsistent across the paper (§3.1 says 508 videos, Table 1 says 100 real/400 fake, §6.1 says 500 videos), the data are only available 'on reasonable request' rather than publicly released, and the paper does not address consent/privacy for the real people whose YouTube videos were manipulated. The evaluation confound is the central weakness; if a matched English baseline were run and the gap persisted, the core claim would be substantially strengthened. Until then, the paper should be treated as a preliminary dataset description whose headline difficulty and Hindi-specificity claims require verification.","tokens_in":86,"tokens_out":6212,"duration_ms":121371,"concrete_test":"Construct a matched English control dataset using the same manipulation tools (FSGAN/FaceSwap/DeepFaceLab for face swap, ReTalking for lip-sync, RVC for voice cloning), the same 15-second/1080p/30fps format, the same numbers of real/fake videos, and the same 408/100 train/test split. Train the same detectors (HeadPose, Two-stream, Mesonet, Xception-c23, Multi-task, FWA, Xception-c40) under identical training conditions and report accuracy with confidence intervals over repeated splits. If HAV-DF accuracy remains substantially below this matched English control, the language-specific explanation gains support; if the gap shrinks or disappears, the lower accuracy is an artifact of pipeline, dataset size, or protocol rather than Hindi content.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that HAV-DF is a Hindi-specific benchmark that is harder to detect rests on the cross-dataset comparison in §5.2.2 (Fig. 13, Tables 3–4). That comparison is uncontrolled. HAV-DF accuracies come from a 408-video training split and 100-video test split (§5.2), while the FF-DF and DFDC figures are drawn from datasets with orders of magnitude more training data (FF++: 1000 real/4000 fake; DFDC: 23,654 real/104,500 fake per Table 1), different generation pipelines, different resolutions and compression, and likely different detector training protocols. The paper does not specify whether the detectors were trained from scratch on HAV-DF, fine-tuned, or applied as pre-trained models, so the reported gap could be explained by train-set size, domain shift from English-trained models, the face-positioning constraint admitted in §6.1, or the specific manipulation toolchain (FSGAN/FaceSwap/DeepFaceLab, ReTalking, RVC) rather than by Hindi-language content. The phrase 'possibly due to its focus on Hindi language content' is therefore an interpretation, not a tested hypothesis. A matched English baseline is required before the dataset can be credited with demonstrating language-driven detection difficulty.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HAV-DF, a Hindi-language audio-video deepfake dataset containing real and manipulated videos generated by face-swapping, lip-sync, and voice cloning. It reports benchmark accuracies of several deepfake detectors on HAV-DF and compares them with published accuracies on English-centric datasets, concluding that HAV-DF is harder to detect and that this is possibly due to Hindi-language content. The paper also claims HAV-DF is the first audio-video deepfake dataset of its kind and includes qualitative samples, SSIM analysis, and a discussion of limitations.","tokens_in":19517,"tokens_out":5228,"duration_ms":40962,"significance":"If the dataset is publicly released and the comparison claims are properly controlled, HAV-DF would fill a genuine gap: no audio-video deepfake dataset exists specifically for Hindi, a major low-resource language with large at-risk populations. The paper's strengths include a reasonable generation pipeline, explicit selection criteria, and an attempt to benchmark multiple detectors. However, the quantitative claim that Hindi language causes lower detection accuracy is not established by the current experiments because the cross-dataset comparison lacks a matched English baseline and does not control for generation pipeline, dataset size, or evaluation protocol. The dataset counts are also internally inconsistent, and the data are not publicly available, limiting reproducibility.","major_comments":[{"comment":"The paper reports inconsistent dataset sizes: §3.1 says 508 videos (200 real + 308 fake); §4.1.2 lists 308 fake (78 Fa+Rv, 130 Ra+Fv, 100 Fa+Fv); §5.2 says 408 training / 100 test (total 508); but §6 and the Limitations say 500 videos, Table 1 reports 100/400, and §3.4.4 says '350 videos' were generated. These inconsistencies make it impossible to understand the actual dataset composition and must be corrected before the dataset can be used as a benchmark.","section":"§3.1, §4.1.2, §5.2, Table 1, §6"},{"comment":"The central claim that HAV-DF is harder to detect than FF-DF and DFDC is based on an uncontrolled cross-dataset comparison. The paper does not state whether detectors were trained from scratch on HAV-DF, fine-tuned, or applied as pre-trained models, nor does it account for the order-of-magnitude differences in training set size (HAV-DF: 408 training videos; FF++: 1000 real/4000 fake; DFDC: 23,654 real/104,500 fake), different generation pipelines, resolutions, compression levels, and face-positioning constraints (admitted in §6.1). No error bars or statistical tests are given. The phrase 'possibly due to its focus on Hindi language content' is therefore an interpretation, not a tested hypothesis. A matched English baseline using the same generation tools and the same evaluation protocol is required to support the language-specific difficulty claim.","section":"§5.2.2, Fig. 13, Tables 3–4"},{"comment":"The abstract states that HAV-DF is 'unique of its kind as all of the previous datasets contain either deepfake videos or synthesized audio,' but the paper itself reviews FakeAVCeleb (§2.3), an audio-video multimodal deepfake dataset, and DFDC, which contains cloned audio and deepfake videos. The novelty claim should be narrowed to 'first Hindi-language audio-video deepfake dataset' or qualified accordingly.","section":"Abstract and §2.3"},{"comment":"The dataset and code are only 'available from the corresponding author on reasonable request.' For a dataset paper, this severely limits reproducibility and community adoption, and it makes the reported benchmark numbers impossible to verify independently. A public release (e.g., a permanent repository with a DOI) or at least an anonymous access mechanism during review is needed for the contribution to be usable as a benchmark.","section":"Data Availability Statement"}],"minor_comments":[{"comment":"The manuscript contains several typographical issues: 'accuracy's' should be 'accuracies' (abstract, §5.2.2), 'multimodel' should be 'multimodal' (e.g., §6), and 'annd'/'hannd' should be 'and'/'hand' (§3.4.3, §4.1.2).","section":"Throughout"},{"comment":"The statement 'Generating high-quality deepfakes for 350 videos' conflicts with the 308 fake video count reported in §4.1.2; this appears to be a typo and should be corrected.","section":"§3.4.4"},{"comment":"The dataset abbreviation 'DFD' used in Table 3 is inconsistent with 'DFDC' used in Figure 12 and Table 4; please unify the terminology.","section":"§5.2.1 and Table 3"},{"comment":"The figures would benefit from error bars or confidence intervals, as the reported differences of a few percentage points are not shown to be statistically significant.","section":"Figures 12 and 13"},{"comment":"The paper refers to 'FF-DF' in §5.2.2 but 'FF++' in earlier sections; please use a consistent abbreviation throughout.","section":"§5.2.2"},{"comment":"Research objective 2 in §1.3 mentions comparing performance on audio-video, audio-only, and video-only, but no audio-only evaluation is reported; either add such experiments or revise the objectives to match the actual scope.","section":"§1.3 and §5"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially valuable if released and the evaluation is properly controlled. The main weaknesses are the uncontrolled cross-dataset comparison and the internal count inconsistencies, both of which are fixable in revision. I would encourage the editor to require a public dataset release as a condition for acceptance, given the paper's nature as a dataset contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a plausible dataset paper with a real gap-filling idea, but the paper as submitted doesn't yet support its headline claims, and the dataset isn't public. Worth one round of review if the authors fix the obvious inconsistencies and release the data.\n\nWhat's actually new: a Hindi-specific audio-video deepfake dataset built with face swap, lip sync, and voice cloning. That combination in a low-resource language is a genuine niche; previous multimodal datasets like FakeAVCeleb, PolyGlotFake, and DFDC don't specifically target Hindi. The authors also document their generation pipeline in detail and include qualitative frame comparisons, SSIM scores, and pitch/spectrogram analyses. The limitations section is honest about the small size and the frontal-view constraint.\n\nThe soft spots are hard to ignore. The video count is inconsistent across the paper: 508, 500, and 100 real/400 fake in Table 1. Section 5.2 says 408 training/100 test, which doesn't obviously match either number. These inconsistencies undermine trust in the artifact. More importantly, the cross-dataset accuracy comparison (Section 5.2.2, Figure 13, Tables 3-4) is uncontrolled. HAV-DF results come from a small training set and a 100-video test set, while FF-DF and DFDC numbers are from datasets orders of magnitude larger, with different generation pipelines and likely different detector training protocols. The paper doesn't say whether detectors were trained from scratch, fine-tuned, or just applied pre-trained. So the lower accuracy on HAV-DF could be due to training-set size, domain shift, or the specific tools (FSGAN, ReTalking, RVC), not the Hindi language per se. The phrase 'possibly due to its focus on Hindi' is an interpretation, not a tested hypothesis. A matched English-language baseline with the same pipeline and protocol is required.\n\nAlso missing: no error bars, no public dataset (only 'available on reasonable request'), and no discussion of consent or privacy for the real YouTube videos, which is an ethical flag for a deepfake dataset.\n\nThe paper would benefit from a serious referee, but only with the expectation of major revision: fix the count inconsistencies, release the dataset with documentation, add a matched baseline, and report confidence intervals. As it stands, I'd treat it as a promising dataset announcement rather than a validated benchmark.\n\nRecommendation: send to peer review, but push hard on data availability and the language-difficulty claim.","headline":"A promising Hindi deepfake dataset paper, but the headline language-difficulty claim is under-supported and the data isn't public yet.","tokens_in":20054,"tokens_out":2016,"would_cite":false,"duration_ms":18210,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HAV-DF, a new Hindi audio-video deepfake dataset, yields lower detection accuracy across seven classifiers than the English-language FF-DF and DFDC benchmarks.","keywords":["deepfake","Hindi dataset","face swap","lip-sync","voice cloning","audio-video deepfake detection","low-resource languages","multimodal deepfake"],"falsifier":"Run the same seven detectors on a matched English-language deepfake corpus generated with the same pipeline (FSGAN, ReTalking, RVC) and matched in resolution, duration, frame rate, video count, and training/test split. If the English corpus yields accuracy comparable to HAV-DF, the central claim that Hindi content drives the difficulty is falsified.","tokens_in":19083,"feed_emoji":"🎭","tokens_out":10537,"duration_ms":76441,"temperature":0.7,"pith_summary":"The paper constructs HAV-DF, a Hindi-language audio-video deepfake dataset, and argues it is the first of its kind: earlier deepfake corpora are either video-only or audio-only, while HAV-DF manipulates both channels. The dataset combines face swap, lip-sync, and voice cloning to create 200 real and 308 fake 15-second videos. The authors benchmark seven existing deepfake detectors on HAV-DF and report lower detection accuracies than on the English-language FF-DF and DFDC datasets. On the paper's telling, a Hindi-specific audiovisual corpus exposes blind spots in detectors built or tuned for English content and provides a benchmark for multilingual deepfake research.","feed_headline":"Hindi deepfake dataset lowers accuracy across seven detectors","feed_subtitle":"First Hindi audio-video deepfake corpus, built with face swap, lip-sync and voice cloning, stumps seven detectors.","key_machinery":"The carrying mechanism is the three-stage generation pipeline: face swap using FSGAN, FaceSwap, and DeepFaceLab; lip-sync using ReTalking to align mouth movements with the audio track; and voice cloning using the RVC model to produce manipulated Hindi speech. The dataset is organized into four sets: real audio with real video ($R_a+R_v$), fake audio with real video ($F_a+R_v$, 78 videos), real audio with fake video ($R_a+F_v$, 130 videos), and fake audio with fake video ($F_a+F_v$, 100 videos), which together allow both unimodal and multimodal deepfake evaluation.","core_discovery":"On its own terms, the paper's claim is that HAV-DF is a new benchmark resource that is harder for current detectors than existing English-centric datasets. On detection methods such as HeadPose, Two-stream, Mesonet, Xception-c23, Multi-task, FWA, and Xception-c40, HAV-DF attains the lowest accuracy among the compared audio-video datasets, with its best result at 64.7% using Xception-c40 versus 99.7% for FF-DF using Xception-c23. The authors attribute this gap to the Hindi-language focus and to the combination of face-swapped, lip-synced, and voice-cloned manipulations, and they position the dataset as the first Hindi-language audio-video deepfake corpus, filling a gap left by video-only and audio-only datasets.","pith_inferences":["A direct test of the language hypothesis would generate an English-language set with the same FSGAN, ReTalking, and RVC pipeline, matched in resolution, duration, frame rate, and video count; if accuracies match HAV-DF, the gap comes from the generation pipeline rather than from Hindi itself.","If Hindi-specific phonetics and gestures drive the difficulty, detectors fine-tuned on HAV-DF should transfer more strongly to other Indic languages than to typologically distant languages, a claim the paper does not test.","The dataset's three manipulation splits support an ablation the paper leaves implicit: measuring how much of the accuracy drop comes from audio cloning, video swapping, and their combination."],"forward_implications":["Existing detectors evaluated on HAV-DF score below their performance on FF-DF and DFDC, indicating that the benchmark is a harder test set for current English-oriented methods.","HAV-DF supplies a resource for training and evaluating models that must detect deepfakes in a low-resource language with audio and video manipulated simultaneously.","Because the fake videos include both celebrity and non-celebrity Hindi speakers across news, technology, and podcast content, models that work on HAV-DF can be tested for generalization across varied Hindi speech and appearance.","The three manipulation combinations ($F_a+R_v$, $R_a+F_v$, $F_a+F_v$) allow a detector to be assessed separately on audio-only, video-only, and joint audio-video forgery.","The reported high SSIM scores indicate that most generated videos are visually close to their real counterparts, so the lower detection accuracy is not simply an artifact of obvious visual flaws."],"supporting_citations":[{"why":"Defines FaceForensics++ (FF-DF), the English-language video dataset that HAV-DF is benchmarked against for detection difficulty.","marker":"Rossler et al., 2019"},{"why":"Defines DFDC, the multimodal deepfake dataset whose detection accuracies are compared with HAV-DF.","marker":"Dolhansky et al., 2020"},{"why":"Introduces FakeAVCeleb, the prior audio-video multimodal deepfake dataset whose language coverage HAV-DF extends.","marker":"Khalid et al., 2021"},{"why":"Provides FSGAN, one of the face-swap generators used to create the manipulated video content.","marker":"Nirkin et al., 2019"},{"why":"Provides ReTalking, the lip-sync method that synchronizes manipulated audio with the video track.","marker":"Cheng, Cun, Zhang, Xia, Yin, Zhu, Wang, Wang and Wang, 2022"},{"why":"Provides the realistic voice cloning (RVC) model used to synthesize fake Hindi speech.","marker":"Kambali, Ansari, Srivastav, Aryan and Nanda, 2023"},{"why":"Supplies HeadPose, one of the seven detection methods whose accuracy on HAV-DF is measured.","marker":"Yang et al., 2019"},{"why":"Supplies Mesonet, a detection model evaluated on HAV-DF.","marker":"Afchar et al., 2018"},{"why":"Supplies the Xception architecture used as Xception-c23 and Xception-c40 in the benchmark.","marker":"Chollet, 2017"},{"why":"Supplies the Two-stream detection method used in the comparative evaluation.","marker":"Zhou et al., 2017"}],"fun_headline_variants":["First Hindi audio-video deepfake dataset stumps detectors","Hindi deepfake corpus combines face, lip, and voice swaps","Hindi deepfakes prove tougher for detectors than English","New Hindi dataset ups the challenge for deepfake detectors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that Hindi makes deepfakes harder to detect rests on the assumption that the lower accuracies on HAV-DF are caused by the language and the diversity of manipulations, not by differences in generation tools, video count, resolution, compression, or evaluation protocol.","fun_headline_variants_meta":{"raw":{"variants":["First Hindi audio-video deepfake dataset stumps detectors","Hindi deepfake corpus combines face, lip, and voice swaps","Hindi deepfakes prove tougher for detectors than English","New Hindi dataset ups the challenge for deepfake detectors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000899,"raw_usage":{"total_tokens":3927,"prompt_tokens":1055,"completion_tokens":2872,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":2806}},"tokens_in":671,"tokens_out":2872,"duration_ms":18806,"temperature":1.0,"reasoning_tokens":2806,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:17:39.541483+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same seven detectors on a matched English-language deepfake corpus generated with the same pipeline (FSGAN, ReTalking, RVC) and matched in resolution, duration, frame rate, video count, and training/test split. If the English corpus yields accuracy comparable to HAV-DF, the central claim that Hindi content drives the difficulty is falsified.","supporting_citations":[{"cited_title":"Xception: Deeplearningwithdepthwiseseparableconvolutions,in: ProceedingsoftheIEEEconferenceoncomputervisionand pattern recognition, pp","cited_arxiv_id":null,"evidence_quote":"Supplies the Xception architecture used as Xception-c23 and Xception-c40 in the benchmark."}],"review_version":1}