{"id":"821cba9e-585e-42f0-b1ce-8c9a7368925e","arxiv_id":"2506.07363","paper_version":2,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A hands-on demonstration that realistic deepfakes can be made for under $160 using Runway, Rope, and ElevenLabs, with a review of the fraud and disinformation risks.","lead":"This student white paper shows that convincing face-swap and voice-clone deepfakes can be produced for about $154 CAD using commercial and open-source tools. It argues that cheap, easy deepfake tools are eroding trust in video, audio, and photo evidence.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central cost-and-realism claim rests on self-assessed output quality; no perceptual or detection validation is provided, so the assertion that realistic deepfakes are producible for under $160 CAD is empirically unsupported.","rationale":"The paper is best read as a demonstration and policy artifact rather than a controlled experiment. Its strongest claim, however, is an empirical statement about cost, accessibility, and realism, and the realism component is the hinge. The reader's UNVERDICTED verdict already reflects this: there is no falsifiable technical result that can be accepted or rejected. My stress-test does not find an internal inconsistency or fraud; it finds that the central empirical assertion is supported only by subjective self-assessment and static figures, and in one place (Section VI.C) the authors explicitly document a visible artifact (hair) that they needed advanced generative AI to address. This does not by itself disprove the claim, but it means the paper has not demonstrated the 'realistic and believable' threshold its conclusion requires. A perceptual and detection test would settle it. I therefore keep the reader's UNVERDICTED verdict unchanged.","tokens_in":15116,"tokens_out":3525,"duration_ms":33166,"concrete_test":"Run a blind study in which at least 50 participants watch the authors' actual deepfake video (or an excerpt) alongside real video of Claudiu Popa and the 2019 Lee/Zuckerberg deepfake, and are asked which are fake and how confident they are; separately run the same media through a state-of-the-art deepfake detector (e.g., a DFDC challenge model for video and a published voice-cloning detector for audio). If participants perform near chance or the detector fails, the realism claim is supported; if participants or the detector reliably flag the authors' media, the 'realistic for under $160' claim fails as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that the authors' self-produced deepfake video and cloned voice are realistic enough to deceive typical viewers and evade detection. The paper offers no evidence of this. Section V.B describes testing restorer models but evaluates only render time, not perceptual fidelity; Section V.A asserts 'high-quality results' via Fig. 5.1, a static still. Section VII concludes the media are 'realistic and believable' and claims GANs make deepfakes 'undecipherable to human eyes and even to an AI model' as a general property, but that is an unsupported non-sequitur, not a measurement of their outputs. The paper's own Section VI.C admits hair movement remained a 'persistent issue' producing 'noticeable imperfections' and required more advanced tools, undercutting the realism claim. The 2019 comparison in Section II.F is also apples-to-oranges: Lee's $772 CAD deepfake (Fig 2.2) had no voice cloning, different software, and a different target; 'superior-looking' is subjective. Because the central claim is that commoditization lets anyone make convincing deepfakes for under $160, the absence of any perceptual study, detection test, or independent review leaves the headline empirical assertion unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This white paper claims that the commoditization of generative AI has sharply lowered the cost and technical skill required to create realistic deepfakes. The authors use Runway, Rope, and ElevenLabs to produce a deepfake video and cloned voice of their project sponsor, reporting a total software cost of $153.61 CAD (about $160 CAD) and comparing their output to a 2019 deepfake by Timothy B. Lee that cost $772 CAD. The paper also surveys deepfake audio tools, evaluates the Rope face-swapping software's restorer models, discusses generative AI models such as GANs and VAEs, and proposes regulatory and technical mitigations including blockchain-based content authentication. The central assertion is that realistic, deception-capable synthetic media are now accessible to anyone with an internet connection for under $160 CAD.","tokens_in":15472,"tokens_out":3200,"duration_ms":29148,"significance":"If the central empirical claim were adequately supported, the paper would have practical value as a demonstration of lowered barriers to deepfake creation, contributing to policy discussions on digital trust. The cost figures, render-time tables, and workflow description are concrete and reproducible in a way that is useful for security awareness. However, the paper's significance is limited by the absence of any perceptual or automated detection evaluation: the claim that the produced media are 'realistic and believable' rests entirely on the authors' own visual judgment and still images. The comparison to the 2019 deepfake is not apples-to-apples, and the paper overgeneralizes from a single demonstration to the broad conclusion that GAN-generated deepfakes are 'undecipherable to human eyes and even to an AI model.' With proper validation and tempered claims, the contribution could be a useful case study, but in its current form the headline claim is empirically underdetermined.","major_comments":[{"comment":"The central claim that a 'superior-looking product' was produced for 'less than a third of the deepfake's price' relies on an invalid comparison. Lee's 2019 deepfake used different software (Face Swap), required manual training on a large dataset, involved no voice cloning, and targeted a different subject (Mark Zuckerberg) with abundant public reference media. The authors' 2024 deepfake targets a sponsor with much less public data, but also uses a different pipeline (Runway for base generation, Rope for face-swapping, ElevenLabs for voice). 'Superior-looking' is asserted subjectively with no perceptual rating or blind comparison. This undermines the quantitative cost comparison in Section II.F and the conclusion in Section V.D that open-source tools are 'child's play' to use.","section":"II.F, V, VII"},{"comment":"The paper provides no evidence that the produced deepfake videos and cloned voices are realistic enough to deceive human viewers or evade detection. Section V.A asserts 'high-quality results' based on a static still (Fig. 5.1), and Section V.B evaluates only render time, not perceptual fidelity or detection robustness. Section VI.C explicitly admits that hair movement remained a 'persistent issue' producing 'noticeable imperfections' that required adoption of more advanced tools. Section VII then concludes that the authors were 'able to create realistic and believable deepfakes' and further claims that GANs make deepfakes 'undecipherable to human eyes and even to an AI model.' This last claim is an unsupported overgeneralization contradicted by a large literature on GAN artifact detection. Without a human perceptual study, an automated detection test, or independent third-party assessment, the paper's core threat-assessment premise is unverified.","section":"V, VI.C, VII"},{"comment":"The render-time comparisons in Tables II and III are reported as single numbers with no indication of variance, number of runs, or experimental protocol. The text does not specify the video resolution, frame count, whether the system was warm, or whether timings include file I/O. Single-run timings on one machine are not sufficient to support the claim of 'performance improvements of up to 4.6 times' or to generalize about the software's resource requirements. This is a presentation and methodology issue, but it affects the paper's practical recommendations about low-cost hardware.","section":"V.B, Tables II and III"},{"comment":"The cost accounting is incomplete and loosely specified. Table I lists final costs of $135.87 for Runway and $17.74 for ElevenLabs but does not state the subscription period, whether taxes or currency conversion are included, or whether the abandoned Amazon EC2 instance (mentioned in V.C.2) incurred any cost. The text moves from '$153.61 CAD' to 'less than $160 CAD' without explaining the discrepancy. Section III presents hypothetical scam scenarios as if they follow from the demonstration, but the paper does not show that real-time voice cloning or real-time video deepfakes were actually implemented, so these scenarios are speculative rather than empirical contributions.","section":"II.E, II.F, III"},{"comment":"The proposed mitigation using blockchain technology is described inconsistently: the text calls it a 'centralized system' after summarizing blockchain as a distributed, tamper-proof ledger. More importantly, the paper cites Gambin et al. and Fraga-Lamas and Fernandez-Carames without critically discussing the known scalability, latency, and adoption barriers of blockchain-based content authentication. As a policy recommendation, this section is too superficial to be actionable.","section":"VII"}],"minor_comments":[{"comment":"The heading 'Obeservations' is a typo for 'Observations', and Section V.B's heading 'Testing & Evaluvation' should be 'Testing & Evaluation'.","section":"IV.B.1"},{"comment":"There is a duplicated word: 'The ElevenLabs companycompany claims' should read 'The ElevenLabs company claims'.","section":"II.E"},{"comment":"The heading 'Mitigation Stratigies' should be 'Mitigation Strategies'.","section":"IV.D.3"},{"comment":"Table I's title 'SOFWARE COSTS (CAD)' contains a typo; it should be 'SOFTWARE COSTS (CAD)'.","section":"II.F"},{"comment":"Figure numbering is inconsistent: Fig. 2.2 is captioned as 'Mark Zuckerberg ... and a deepfake ...' while the text refers to 'Fig. 2.3', and Section V.C.1 refers to 'Fig. 3.2' which does not exist in the manuscript. The name 'Timothee B. Lee' in the Fig. 2.2 caption should be 'Timothy B. Lee'.","section":"II.F and V"},{"comment":"The sentence about Fig. 5.1 is ambiguous: it says 'On the left, we see Rupert Friend as Agent 47 ... and on the right is Claudiu Popa ... face swapped into the clip'—clarify which side is the original and which is the swap.","section":"V.A"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads as a course project or white paper rather than a peer-reviewed research article. Its main value is as a practical demonstration of tool accessibility, but the authors need to either add rigorous evaluation (perceptual study, detection tool tests, independent review) or substantially weaken the realism and undetectability claims. The 2019 cost comparison should be reframed as anecdotal rather than a controlled comparison. If the venue accepts practitioner-oriented case studies, a revised version with stripped-down claims and explicit limitations could be a valid contribution; otherwise, the current version does not meet the empirical bar of a research paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a student white paper, not a research preprint, and the reader's take is roughly right: it re-demonstrates a known deepfake pipeline and packages it as a policy warning. The genuinely new element is a narrow cost data point: the authors produced a ~4x longer video with voice cloning for $153.61 CAD using Runway, Rope (or Rope NEXT), and ElevenLabs, compared with Timothy Lee's $772 CAD face-only deepfake in 2019. They also report useful render-time comparisons across Rope restorers and show Rope NEXT's TensorRT fork gives up to a 4.6x speedup on an RTX 3060 laptop. Those numbers are concrete and reproducible from their stated hardware and software choices.\n\nWhat the paper does well is document the actual workflow: tool choices, costs, settings, and honest descriptions of failures (hair movement, VRAM limits, EC2 installation pain). The literature review is broad and appropriately contextualizes the demonstration in known fraud cases and mitigation efforts. The writing is clear for a course project.\n\nThe soft spots are real and proportionately important. The central claim—that realistic deepfakes are now a sub-$160 commodity—rests on the authors' own judgment of their outputs. There is no perceptual study, no detection experiment, no independent evaluation. The paper's own Section VI.C admits hair movement remained a 'persistent issue' with 'noticeable imperfections,' which undercuts the Section VII assertion that the media were 'realistic and believable.' The Section V.D claim that GANs make deepfakes 'undecipherable to human eyes and even to an AI model' is an unsupported general statement, not a measurement of their outputs. The 2019 comparison is also apples-to-oranges: different software, different target, no voice cloning, and subjectively assessed 'superior-looking.' So the headline empirical assertion is plausible but unverified.\n\nWho is this for? A reader wanting a concrete, current example of the cost and effort of deepfake production for a policy or awareness discussion could get some use out of it. A researcher looking for a falsifiable claim or rigorous evaluation will not find one.\n\nRecommendation: desk reject for peer review. It is not a research contribution; it is a competent student report. If it were submitted to a venue with a demonstration track or industry report track, a reviewer could ask for a perceptual/detection validation before accepting. As an arXiv preprint, it is harmless and may be cited as anecdotal evidence, but it does not warrant referee time as a research paper.","headline":"A competent student white paper that re-demonstrates a known deepfake workflow with a concrete but unvalidated cost data point; not a research contribution.","tokens_in":15793,"tokens_out":2489,"would_cite":false,"duration_ms":25887,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Realistic deepfakes are now a commodity: two subscriptions and free open-source tools produced a face-swapped, voice-cloned video for $153.61 CAD, less than a third of the 2019 cost.","keywords":["deepfakes","generative AI","voice cloning","face-swapping","media authenticity","digital trust","deepfake fraud","cost analysis"],"falsifier":"If untrained human raters, or off-the-shelf deepfake detectors, consistently identify the authors' videos and cloned voices as fake, the central claim that convincing deepfakes are now available for under $160 CAD would be undercut. A simpler check is reproduction: a different team following the same workflow with the same budget and hardware would need to obtain comparable quality for the generalization to hold.","tokens_in":14955,"feed_emoji":"🎭","tokens_out":7805,"duration_ms":69123,"temperature":0.7,"pith_summary":"This white paper argues that realistic deepfakes have become a commodity. With two paid subscriptions and free open-source software, the authors produced a face-swapped, voice-cloned video they judge better than a comparable 2019 deepfake that cost about $772 CAD. The paper's goal is to show that the financial and technical barriers to creating convincing synthetic media have collapsed, and that this collapse threatens trust in video and audio evidence. It also evaluates voice-cloning tools, tests face-swap restorers, and calls for regulation, detection, and public awareness. If the central demonstration holds, the dominant threat is not specialized AI labs but ordinary fraudsters with an internet connection.","feed_headline":"A $160 deepfake now beats a 2019 $772 deepfake","feed_subtitle":"The authors cloned a CEO's face and voice for $153.61 CAD using off-the-shelf tools.","key_machinery":"The central mechanism is a three-tool production workflow. Runway's Gen-3 Alpha generates the base video and provides lip-syncing for the scripted face; Rope performs the face swap, accepting a single still image of the target and offering restorer models (GPEN256, GFPGAN, CF, GPEN512) that sharpen the swapped face; ElevenLabs' Creator plan clones the target's voice and can change the actor's voice into the victim's for live scenarios. The quantitative core of the argument is the cost comparison: $153.61 CAD for the 2024 pipeline versus $772 CAD for the 2019 baseline, with side-by-side stills used to assert the newer output looks more realistic.","core_discovery":"The paper's central claim is that the commoditization of generative AI lets a non-specialist produce a deepfake that is superior to the 2019 state of the art for less than a third of the cost. Using Runway's Gen-3 Alpha to generate and lip-sync base video, Rope (an open-source fork of Roop) to swap faces from a single reference image, and ElevenLabs to clone a voice from short audio samples, the authors produced a longer, smoother video with realistic voice cloning for $153.61 CAD. Their baseline is a 2019 deepfake that cost $552 USD ($772 CAD), was visual-only, and looked jittery and unnatural. From this comparison they conclude that realistic deepfakes are now available 'to anyone with an internet connection,' making the erosion of digital trust an immediate practical problem rather than a hypothetical one.","pith_inferences":["A formal perceptual study asking untrained viewers to distinguish the authors' deepfakes from real footage would quantify the threat; the paper stops at side-by-side figures and subjective wording.","The $153.61 CAD figure counts subscriptions only, not labour time, hardware, or skill; including those would raise the effective barrier and qualify the 'anyone with an internet connection' claim.","The same pipeline could be benchmarked against commercial deepfake detectors to identify which artifacts, such as hair motion, lip-sync lag, or restorer choice, are most detectable; that experiment is not run here.","The paper's appeal to blockchain-style authentication implies a testable requirement: provenance systems must keep verification costs lower than the cost of producing a convincing fake, an economic constraint the paper leaves implicit."],"forward_implications":["A fraudster with no AI expertise can impersonate a specific person using one still image and short audio samples, because the entire pipeline runs on consumer hardware.","Live impersonation on video calls is within reach, since Rope runs on a 6 GB laptop GPU and ElevenLabs supports real-time voice changing.","Audio-only scams can be automated and scaled, because cloned voices can answer phone calls in real time and be generated in bulk.","Video and audio lose their default status as trustworthy evidence, pushing courts, insurers, and platforms toward authentication or detection.","Detection is a moving target, because new forks of the face-swap software keep improving rendering speed and quality, as the Rope NEXT speedups show."],"supporting_citations":[{"why":"Supplies the 2019 baseline: a $552 USD visual-only deepfake that the paper's own $153.61 CAD production is compared against.","marker":"[11]"},{"why":"Documents Runway and its Gen-3 Alpha Unlimited plan, the base-video generation and lip-sync tool that dominates the cost.","marker":"[6]"},{"why":"Documents ElevenLabs' voice cloning and Creator plan used to produce the cloned voice.","marker":"[9]"},{"why":"Defines Roop, the original open-source one-click face-swap program on which Rope is built.","marker":"[7]"},{"why":"Describes Rope's GUI, restorer options, and cosmetic correction tools used in the face-swap pipeline.","marker":"[8]"},{"why":"Provides the $35 million voice-cloning bank fraud example used to motivate the threat.","marker":"[3]"}],"fun_headline_variants":["A $153 deepfake outshines 2019's $772 effort","Deepfake quality leaps as price drops to $153","From $772 to $153: Better deepfakes for less","The $153 deepfake that beats a $772 one","Affordable AI: $153 deepfake rivals 2019's $772"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the authors' self-produced videos and cloned voices are realistic enough to deceive typical viewers and listeners; the paper supports this with side-by-side figures and subjective wording, not with a perception study or detection test.","fun_headline_variants_meta":{"raw":{"variants":["A $153 deepfake outshines 2019's $772 effort","Deepfake quality leaps as price drops to $153","From $772 to $153: Better deepfakes for less","The $153 deepfake that beats a $772 one","Affordable AI: $153 deepfake rivals 2019's $772"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000335,"raw_usage":{"total_tokens":1827,"prompt_tokens":883,"completion_tokens":944,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":854}},"tokens_in":499,"tokens_out":944,"duration_ms":8767,"temperature":1.0,"reasoning_tokens":854,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:53:26.302669+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If untrained human raters, or off-the-shelf deepfake detectors, consistently identify the authors' videos and cloned voices as fake, the central claim that convincing deepfakes are now available for under $160 CAD would be undercut. A simpler check is reproduction: a different team following the same workflow with the same budget and hardware would need to obtain comparable quality for the generalization to hold.","supporting_citations":[{"cited_title":"I created my own deepfake -it took two weeks and cost $552,","cited_arxiv_id":null,"evidence_quote":"Supplies the 2019 baseline: a $552 USD visual-only deepfake that the paper's own $153.61 CAD production is compared against."},{"cited_title":"Tools for human imagination.,","cited_arxiv_id":null,"evidence_quote":"Documents Runway and its Gen-3 Alpha Unlimited plan, the base-video generation and lip-sync tool that dominates the cost."},{"cited_title":"Free text to Speech & AI Voice Generator,","cited_arxiv_id":null,"evidence_quote":"Documents ElevenLabs' voice cloning and Creator plan used to produce the cloned voice."},{"cited_title":"S0MD3V/Roop: One -click face swap,","cited_arxiv_id":null,"evidence_quote":"Defines Roop, the original open-source one-click face-swap program on which Rope is built."},{"cited_title":"Gui Ruby Rope Face Swap,","cited_arxiv_id":null,"evidence_quote":"Describes Rope's GUI, restorer options, and cosmetic correction tools used in the face-swap pipeline."},{"cited_title":"Deepfaked voice enabled $35 million bank heist in 2020,","cited_arxiv_id":null,"evidence_quote":"Provides the $35 million voice-cloning bank fraud example used to motivate the threat."}],"review_version":1}