Pith. sign in

REVIEW 3 major objections 5 minor 112 references

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read BounTCHA claims a CAPTCHA can rest on people's ability to spot where a real video ends and an AI-generated extension begins, while current multimodal models largely miss it.

desk verdict A genuinely new CAPTCHA idea with a solid human-performance study, but the security analysis ignores the cheapest attack class—classical cut detection—so the resilience claim is not yet established. read the letter →

arxiv 2501.18565 v3 pith:MW3TXQXA submitted 2025-01-30 cs.CR cs.AIcs.HC

classification cs.CRcs.AIcs.HC
keywords CAPTCHAvideoboundaryidentificationgenerativeAIextensionhumanperceptionmultimodalLLMwebsecuritytimebiashuman-botdiscrimination
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

BounTCHA is a proposed CAPTCHA built on a perceptual asymmetry: people quickly notice when a video suddenly shifts into something unnatural, while current AI video-understanding systems do not. The paper claims to be the first to make that asymmetry the entire test, asking users to drag a progress bar to the moment where a real clip ends and an AI-generated extension begins. A guided pipeline produces the material: a video-language model describes the clip, an LLM writes a prompt that pushes the story in an unexpected direction, and a video generator extends the clip, after which the two parts are stitched and compressed into a short web-friendly video. In a study of 186 participants the human time-bias distribution is roughly normal, and the paper reports human success of at least 82% at $\alpha \le 0.25$, while four tested multimodal models (Tarsier, MiniCPM-V 2.6, GPT-4V, Claude 3.5 Sonnet) exceed no more than 17.33% success. If those numbers hold, BounTCHA would give websites a CAPTCHA that current AI-powered bots cannot solve.

What carries the argument

The load-bearing object is the guided AI-extended composite video: a raw real-camera clip concatenated with a prompted AI-generated continuation, with the true seam at frame $F^*_n$ serving as ground truth. The decision mechanism is the human time-bias acceptance window: a submitted boundary time is accepted when it falls in $[\beta_1,\beta_2]$, where $\beta_i = \mu \pm \sigma \cdot \mathrm{ppf}(1-\alpha/2)$ and $\mu, \sigma$ come from the human study; the significance level $\alpha$ and the video length $L$ are the tunable knobs that control how narrow the window is. The generation pipeline carries the argument by making the seam salient to human perception but not to the tested multimodal models. The paper also uses shift cutting to relocate the boundary inside a single generated video, turning one generation into multiple CAPTCHA variants.

What would settle it

Run a standard cut-detection pass, such as frame-by-frame histogram difference or FFmpeg's scene-detection filter, on the 25 released composite videos and check whether the detected seam falls inside the human acceptance window at $\alpha = 0.25$. If a cheap detector locates the boundary with far higher accuracy than the reported MLLM success rates, BounTCHA's security claim is refuted.

Watch

Extended reading notes

Core claim

The central claim is that the boundary between a raw video and a guided AI-generated continuation is itself a usable human test: people can locate it reliably, leading multimodal models cannot, and the resulting distribution of human errors supports a simple pass/fail rule. The paper calls the boundary the 'real boundary frame' $F^*_n$ and treats it as ground truth. It reports a human time-bias distribution with mean $0.332$ s and standard deviation $0.406$ s, and defines the acceptance window from that distribution as $\mu \pm \sigma \cdot \mathrm{ppf}(1-\alpha/2)$. Leave-one-out cross-validation over the 25 videos gives human success rates that stay at or above 82% for $\alpha \le 0.25$; under the same window the four tested MLLMs succeed in at most 17.33% of trials and take over 20 seconds per attempt, against 14.2 seconds average for humans. The paper's conclusion is that a boundary-identification task of this kind can currently separate humans from state-of-the-art multimodal AI.

Load-bearing premise

The security case assumes attackers will only guess randomly, search a database of known videos, or ask a multimodal AI model; it never tests simple, cheap algorithms that look for an abrupt visual cut at the stitching point, and such an algorithm would break the CAPTCHA if it works.

Editorial extensions

If this is right

  • If BounTCHA's numbers hold, generative AI itself becomes a source of CAPTCHA material, and the same models that erode text- and image-based CAPTCHAs can be used to build new challenges.
  • A deployment can set the acceptance window from the reported human distribution and tune $\alpha$ and video length to trade human success against bot success.
  • Current multimodal LLMs would need to reach well above 17.33% boundary-location accuracy before an MLLM-based solver becomes a practical threat to BounTCHA.
  • Because human time-bias statistics are roughly normal and age-group differences are small, one window can serve users across the tested age range.
  • Shift cutting provides a cheap way to generate fresh variants from each produced video, which the database-attack analysis treats as protection against replay.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper never tests cheap, non-semantic seam detectors such as frame-histogram differences or optical-flow discontinuities; my inference is that those, not MLLMs, are the threat most likely to break BounTCHA, because the composite video is made by simple concatenation.
  • A testable extension would be to soften the seam with a crossfade or motion continuity and then re-measure both human and classical-detector performance; if humans stay accurate while detectors degrade, the mechanism becomes much stronger.
  • The same 'boundary between real and AI-generated' design could generalize to other perceptual discontinuities, such as audio-visual mismatches, physics-violating motion, or story-consistency breaks, wherever humans' world knowledge currently outpaces MLLMs.
  • If video-language models keep improving, the acceptance window should be recalibrated over time; the paper's normal-distribution model gives a direct procedure for doing that, but the paper itself does not propose a retraining or update schedule.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes BounTCHA, a CAPTCHA in which a user watches a short video formed by concatenating a raw clip with a generative-AI extension and drags a progress bar to the reported boundary. The generation pipeline uses Tarsier for video description, GPT-4o for prompt synthesis, and Kling for video extension (Section 3). A user study with 186 participants on 25 videos measures the time bias between the true boundary and the user-reported boundary; the acceptance interval is defined in Eq. (4) from the fitted mean and standard deviation of that bias. The paper claims human success rates of at least 82% at significance level alpha <= 0.25, and evaluates uniform and truncated-normal random attacks, a database attack, and MLLM attacks (Tarsier, the model referred to as MiniVPM-V 2.6, GPT-4V, and Claude 3.5 Sonnet), with claimed attack success rates no higher than 17.33%.

Significance. If the security claims held, BounTCHA would be a timely and genuinely novel CAPTCHA: it exploits a perceptual capability (sensitivity to abrupt video transitions) that current MLLMs do not reliably exhibit, and the authors ship an open-sourced prototype, the video set, and a real user study. The arithmetic of the random-attack analysis is straightforward, and the database-attack model is clearly derived. However, the headline numbers are weakened by two load-bearing gaps: the human acceptance threshold is derived from the same data used to report human success, and the security analysis omits cheap classical video-segmentation attacks that are directly aimed at the kind of concatenation boundary the scheme creates. The MLLM evaluation is also too small and under-specified to support a robust upper-bound claim.

major comments (3)
  1. [Section 5.2 and Eq. (4)] The reported human success rates are partly true by construction. The acceptance interval [beta1, beta2] is defined from the fitted mean mu and standard deviation sigma of the same 186-participant time-bias data that are then classified as successes, so under the normal model the aggregate success rate at confidence 1-alpha is approximately 1-alpha; the green curve in Fig. 9 therefore inherits its value from the Gaussian fit rather than from an independent behavioral measurement. The leave-one-out validation divides the 25 videos into five groups but reuses the same participants in every fold, so it does not establish generalization to new users. Please report held-out participant success (participant-level cross-validation or a fresh deployment), fix the threshold before evaluation, and show the distribution of per-participant success instead of only the aggregate curve.
  2. [Section 6 and Section 3.3 (Eq. 3)] The security analysis omits classical video-boundary detection, which is the most obvious low-cost attack. Each challenge video is the direct concatenation of V_in and V_ext, with F*_n as the transition frame; per-frame color-histogram distance, optical-flow discontinuity scoring, and FFmpeg scene detection are designed to locate exactly such abrupt changes and can run in milliseconds. The paper should measure at least one such baseline on the 25 published videos and report how often the detected boundary falls within [beta1, beta2] at alpha = 0.25. Without this measurement, the claim that BounTCHA is resilient against various types of attacks is unsupported for the cheapest attack class.
  3. [Section 6.3 and Table 2] The MLLM attack evaluation is too small and under-specified to support the stated upper bound. With 25 videos and three rounds per model, a 17.33% success rate corresponds to a 95% confidence interval of roughly 9-28%, so the claim that the models succeed in no more than 17.33% of trials is not a robust statistical statement. The paper does not report the prompt template, decoding parameters, or frame-sampling strategy for the four models, which limits reproducibility. Please provide this information, report per-video and per-round breakdowns, and compare against human performance using a non-circular acceptance threshold, ideally at fixed human false-positive rates.
minor comments (5)
  1. [Section 6.3 and Figure 15] The sentence claiming that the oldest group's worst-case time is notably higher than the worst-case time for other MLLMs contradicts Table 2, where the human worst-case time is 26.6s and all MLLM worst-case times are between 37.1s and 43.2s; please correct the text or the figure labels.
  2. [Table 2 and Section 6.3] The model name 'MiniVPM-V 2.6' appears to be a typo for MiniCPM-V 2.6 (reference [100]); please standardize the name throughout.
  3. [Eq. (4)] Writing beta_i = mu ± sigma * ppf(1 - alpha/2) with i in {1,2} is ambiguous; please define beta_1 = mu - sigma * ppf(1 - alpha/2) and beta_2 = mu + sigma * ppf(1 - alpha/2) explicitly.
  4. [Section 3.2.2] The generation prompt says the extension should 'differ significantly' from the original video while also avoiding 'drastic changes in the visuals'; please clarify how these constraints are reconciled, since both directly affect task difficulty.
  5. [Throughout] Minor typographical issues include 'CATPCHA' in Figure 1, 'gameified' in Section 2.1.3, and 'not only happened' in Section 1; a light copyedit would improve readability.

Circularity Check

1 steps flagged · score 4.0 of 10

The human success-rate claim restates the fitted confidence interval; MLLM attack results are external, so circularity is partial.

  1. fitted input called prediction [Section 5.2 (Figure 9), Section 6 Eq. (4), Table 2]
    "All the data of time bias falls between -1.6s and 2.6s, and it follows a normal distribution. Therefore, we can increase the difficulty of BounTCHA by adjusting the significance level to narrow the time range. ... β_i = μ ± σ · ppf(1−α/2), i∈{1,2} (4) ... Human ≥ 82% (α≤0.25), The relevant data is from Figure 9."

    The acceptance interval [β1, β2] in Eq. (4) is defined as the central (1−α) interval of the normal distribution fitted to the same 186-participant time-bias data (μ = 0.332, σ = 0.406). A human drawn from that same distribution therefore falls inside the interval with probability approximately 1−α by construction. Table 2's headline 'Human ≥ 82% (α≤0.25)' is thus a restatement of the fitted interval's coverage, not an independent measurement of human ability. The leave-one-out cross-validation only verifies that the fitted distribution generalizes to held-out videos from the same dataset; it does not break the definitional link between the fitted confidence level and the reported success rate. The MLLM attack results are genuinely external, which keeps the circularity partial.

full rationale

The main derivation chain is not globally circular: the video-generation pipeline, the 186-participant user study, and the MLLM attack benchmarks are independent empirical contributions. The one genuinely circular element is the presentation of the human success rate. Because Eq. (4) defines the acceptance window as a central (1−α) interval of the fitted human time-bias distribution, the success rates plotted in Figure 9 and quoted in Table 2 are close to 1−α by construction; the cross-validation gives limited independence but does not remove the definitional dependence. The self-citation [90] is motivational only and is backed by this paper's own new user study, so it is not load-bearing. The security analysis in Section 6 omits classical shot/cut-detection baselines (e.g., histogram-based or optical-flow boundary detection); that is a security-evaluation gap rather than a circularity and does not affect the circularity score. Overall score 4 reflects one partial by-construction result while the core MLLM-vs-human comparison remains externally measured.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central security and usability claims rest on two fitted parameters (mu, sigma), one policy threshold (alpha), and several unverified assumptions: normality of the time-bias distribution, uniform boundary placement across variants, and the absence of simple video-segmentation attacks. The normality and uniform-boundary assumptions are stated in the paper; the absence of segmentation attacks is an omission that the security analysis does not acknowledge.

free parameters (3)
  • mu = 0.332 s
    Mean time bias of human boundary identifications, fitted to 186 participants x 25 videos; defines the center of the acceptance window in Eq. (4).
  • sigma = 0.406 s
    Standard deviation of human time bias, fitted to the same dataset; controls the width of the acceptance window and the random-attack success probability in Eq. (5).
  • alpha = 0.25 (recommended)
    Significance level chosen by the designers to balance human success (>= 82%) against random-attack success; it is a policy knob rather than a fitted value, but the central security claims depend on it.
assumptions (4)
  • domain assumption Human time bias follows a normal distribution
    Stated in Section 5.2 ('it follows a normal distribution') without a statistical test; the entire acceptance window in Eq. (4) and all derived attack probabilities depend on this shape.
  • domain assumption Variants within a video group have uniformly distributed boundaries
    Assumed in Section 6.2 to obtain gamma_i = 1/(U_i-u_i) and the simplified P = m*u/(M*U) + 1/U; if boundaries cluster, the database attack success differs.
  • ad hoc to paper The concatenation point is not detectable by low-level video segmentation
    Used implicitly by the security analysis in Section 6, which never tests classical cut-detection algorithms; this is the weakest security premise.
  • standard math Trial independence for the binomial attack model
    Used in Eq. (12) for the number of successful attacks in n rounds; standard but unstated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos." pith.science (2026). https://pith.science/paper/MW3TXQXA

@misc{pith2026250118565,
  author       = {Pith},
  title        = {Pith review of: BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MW3TXQXA}},
  note         = {Machine review of arXiv:2501.18565}
}
read the original abstract

In recent years, the rapid development of artificial intelligence (AI) especially multi-modal Large Language Models (MLLMs), has enabled it to understand text, images, videos, and other multimedia data, allowing AI systems to execute various tasks based on human-provided prompts. However, AI-powered bots have increasingly been able to bypass most existing CAPTCHA systems, posing significant security threats to web applications. This makes the design of new CAPTCHA mechanisms an urgent priority. We observe that humans are highly sensitive to shifts and abrupt changes in videos, while current AI systems still struggle to comprehend and respond to such situations effectively. Based on this observation, we design and implement BounTCHA, a CAPTCHA mechanism that leverages human perception of boundaries in video transitions and disruptions. By utilizing generative AI's capability to extend original videos with prompts, we introduce unexpected twists and changes to create a pipeline for generating guided short videos for CAPTCHA purposes. We develop a prototype and conduct experiments to collect data on humans' time biases in boundary identification. This data serves as a basis for distinguishing between human users and bots. Additionally, we perform a detailed security analysis of BounTCHA, demonstrating its resilience against various types of attacks. We hope that BounTCHA will act as a robust defense, safeguarding millions of web applications in the AI-driven era.

Figures

Figures reproduced from arXiv: 2501.18565 by the authors.

Figure 1
Figure 1. Common text-based and image-based CAPTCHA examples, including arithmetic CAPTCHA, reCAPTCHA, puzzle CAPTCHA, [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Showcases of 3D & gamified CAPTCHAs. (a) is a text-based 3D CAPTCHA [ [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The production pipeline for generating BounTCHA videos. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: A bar chart comparing the sizes of original videos ( [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: The time cost of the video preparation pipeline. The length of the blocks is not drawn to scale based on the time duration. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: User interface of the BounTCHA prototype. [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: The procedure of user study [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Left: length of videos and distribution of time bias between the human identification boundary and the actual boundary for [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: The overall time bias range where the confidence levels with the corresponding significance levels. And their overall success [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 10
Figure 10. Figure 10: The age group pie chart of the participants. [PITH_FULL_IMAGE:figures/full_fig_p012_10.png]
Figure 11
Figure 11. Figure 11: Left: the time bias range and the success rate of the group in the age 18-29. Mid: in the age 30-39. Right: in the age 40-48. [PITH_FULL_IMAGE:figures/full_fig_p013_11.png]
Figure 12
Figure 12. Figure 12: 3D Surface Plot of 𝑃 (𝑋 ∈ [𝛽1, 𝛽2 ] ): Visualization of the 𝑃 (𝑋 ∈ [𝛽1, 𝛽2 ] ) in respect to variables 𝐿 and 𝛼. 𝜑(𝜉) = 1 √ 2𝜋 exp(−1 2 𝜉 2 ), Φ(𝑥) = 1 2 (1 + erf(𝑥/ √ 2)) (7) and 𝑓 = 0 otherwise, where erf(·) is the Gauss error function erf(𝑧) = 2 √ 𝜋 ∫ 𝑧 0 exp(−𝑡 2 )…
Figure 13
Figure 13. Figure 13: 3D Surface Plot of truncated normal distribution random attack success probability in respect to the time bias range proportion [PITH_FULL_IMAGE:figures/full_fig_p015_13.png]
Figure 14
Figure 14. Figure 14: Left: Probability density with 𝑀 = 1000 and 𝑈 = 10, where 𝜔1 = 𝑚 𝑀 and 𝜔2 = 𝑢 𝑈 ranging from 0.3 to 0.9. Alongside the number of attacks, it shows the corresponding probabilities. Right: Similar to the left, with 𝑀 = 1000 and 𝑈 = 100. 6.3 Multi-modal LLM (MLLM) Attack…
Figure 15
Figure 15. Figure 15: The chart illustrating the relationship between age groups and the average time spent, the longest time spent, and success [PITH_FULL_IMAGE:figures/full_fig_p017_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

112 extracted references · 59 canonical work pages

  1. [1]

    Suhas Aggarwal. 2013. Animated CAPTCHAs and games for advertising. In Proceedings of the 22nd International Conference on World Wide Web . 1167–1174

  2. [2]

    YeonChan Ahn, Namsoo Kim, and Yoo-Sung Kim. 2013. A user-friendly image-text fusion CAPTCHA for secure web services. In Proceedings of International Conference on Information Integration and Web-based Applications & Services . 550–554

  3. [3]

    Omar Alonso, Catherine Marshall, and Marc Najork. 2013. A human-centered framework for ensuring reliability on crowdsourced labeling tasks. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol. 1. 2–3

  4. [4]

    Omar Alonso, Catherine Marshall, and Marc Najork. 2014. Crowdsourcing a subjective labeling task: a human-centered framework to ensure reliable results. Microsoft Res., Redmond (2014)

  5. [5]

    Fatmah H Alqahtani and Fawaz A Alsulaiman. 2020. Is image-based CAPTCHA secure against attacks based on machine learning? An experimental study. Computers & Security 88 (2020), 101635

  6. [6]

    Babak Amin Azad, Oleksii Starov, Pierre Laperdrix, and Nick Nikiforakis. 2020. Web runner 2049: Evaluating third-party anti-bot services. In Detection of Intrusions and Malware, and Vulnerability Assessment: 17th International Conference, DIMV A 2020, Lisbon, Portugal, June 24–26, 2020, Proceedings 17. Springer, 135–159

  7. [7]

    K Anjitha and IK Rijin. 2015. Captcha as graphical passwords-enhanced with video-based captcha for secure services. In 2015 International Conference on Applied and Theoretical Computing and Communication Technology (iCATccT) . IEEE, 213–217

  8. [8]

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. (2023)

Show all 112 references
  1. [9]

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. 2023. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 (2023)

  2. [10]

    Neha Pradyumna Bora and Dinesh Chandra Jain. 2023. A web authentication biometric 3D animated CAPTCHA system using artificial intelligence and machine learning approach. Journal of Artificial Intelligence and Technology 3, 3 (2023), 126–133

  3. [11]

    Jose Brustoloni. 2002. Protecting electronic commerce from distributed denial-of-service attacks. In Proceedings of the 11th international conference on World Wide Web. 553–561. BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos 19

  4. [12]

    Elie Bursztein, Steven Bethard, Celine Fabry, John C Mitchell, and Dan Jurafsky. 2010. How good are humans at solving CAPTCHAs? A large scale evaluation. In 2010 IEEE symposium on security and privacy . IEEE, 399–413

  5. [13]

    Wenhao Chai, Xun Guo, Gaoang Wang, and Yan Lu. 2023. Stablevideo: Text-driven consistency-aware diffusion video editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 23040–23050

  6. [14]

    Arindam Chaudhuri, Krupa Mandaviya, Pratixa Badelia, Soumya K Ghosh, Arindam Chaudhuri, Krupa Mandaviya, Pratixa Badelia, and Soumya K Ghosh. 2017. Optical character recognition systems . Springer

  7. [15]

    Chaoran Chen, Leyang Li, Luke Cao, Yanfang Ye, Tianshi Li, Yaxing Yao, and Toby Jia-jun Li. 2024. Why am I seeing this: Democratizing End User Auditing for Online Content Recommendations. arXiv preprint arXiv:2410.04917 (2024)

  8. [16]

    Jun Chen, Xiangyang Luo, Yanqing Guo, Yi Zhang, and Daofu Gong. 2017. A Survey on Breaking Technique of Text-Based CAPTCHA. Security and communication networks 2017, 1 (2017), 6898617

  9. [17]

    Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. 2018. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV) . 801–818

  10. [18]

    Monica Chew and J Doug Tygar. 2004. Image recognition captchas. In International Conference on Information Security . Springer, 268–279

  11. [19]

    Joseph Cho, Fachrina Dewi Puspitasari, Sheng Zheng, Jingyao Zheng, Lik-Hang Lee, Tae-Ho Kim, Choong Seon Hong, and Chaoning Zhang. 2024. Sora as an agi world model? a complete survey on text-to-video generation. arXiv preprint arXiv:2403.05131 (2024)

  12. [20]

    Paul Couairon, Clément Rambour, Jean-Emmanuel Haugeard, and Nicolas Thome. 2023. Videdit: Zero-shot and spatially aware text-driven video editing. Transactions on Machine Learning Research (2023)

  13. [21]

    Yifan Cui, Xinyi Shan, and Jeanhun Chung. 2024. A Feasibility Study on RUNWAY GEN-2 for Generating Realistic Style Images. International Journal of Internet, Broadcasting and Communication 16, 1 (2024), 99–105

  14. [22]

    Alexey Dosovitskiy. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)

  15. [23]

    John R Douceur. 2002. The sybil attack. In International workshop on peer-to-peer systems . Springer, 251–260

  16. [24]

    Zhengjie Du, Yuekang Li, Yaowen Zheng, Xiaohan Zhang, Cen Zhang, Yi Liu, Sheikh Mahbub Habib, Xinghua Li, Linzhang Wang, Yang Liu, et al

  17. [25]

    Peter Eckersley. 2010. How unique is your web browser?. In Privacy Enhancing Technologies: 10th International Symposium, PETS 2010, Berlin, Germany, July 21-23, 2010. Proceedings 10 . Springer, 1–18

  18. [26]

    Tuğrulcan Elmas. 2023. Analyzing activity and suspension patterns of twitter bots attacking turkish twitter trends by a longitudinal dataset. In Companion Proceedings of the ACM Web Conference 2023 . 1404–1412

  19. [27]

    Christoph Fritsch, Michael Netter, Andreas Reisser, and Günther Pernul. 2010. Attacking Image Recognition Captcha s: A Naive but Effective Approach. In Trust, Privacy and Security in Digital Business: 7th International Conference, TrustBus 2010, Bilbao, Spain, August 30-31, 20...

  20. [28]

    Haichang Gao, Dan Yao, Honggang Liu, Xiyang Liu, and Liming Wang. 2010. A novel image based CAPTCHA using jigsaw puzzle. In 2010 13th IEEE international conference on computational science and engineering . IEEE, 351–356

  21. [29]

    Paolo Gasti, Gene Tsudik, Ersin Uzun, and Lixia Zhang. 2013. DoS and DDoS in named data networking. In 2013 22nd International Conference on Computer Communication and Networks (ICCCN) . IEEE, 1–7

  22. [30]

    Nethanel Gelernter and Amir Herzberg. 2016. Tell me about yourself: The malicious captcha attack. In Proceedings of the 25th International Conference on World Wide Web. 999–1008

  23. [31]

    Rich Gossweiler, Maryam Kamvar, and Shumeet Baluja. 2009. What’s up CAPTCHA? A CAPTCHA based on image orientation. In Proceedings of the 18th international conference on World wide web . 841–850

  24. [32]

    Gilad Gressel, Rahul Pankajakshan, and Yisroel Mirsky. 2024. Discussion Paper: Exploiting LLMs for Scam Automation: A Looming Threat. In Proceedings of the 3rd ACM Workshop on the Security Implications of Deepfakes and Cheapfakes . 20–24

  25. [33]

    Meriem Guerar, Luca Verderame, Mauro Migliardi, Francesco Palmieri, and Alessio Merlo. 2021. Gotta CAPTCHA’Em all: a survey of 20 Years of the human-or-computer Dilemma. ACM Computing Surveys (CSUR) 54, 9 (2021), 1–33

  26. [34]

    Chris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk, and Philipp Zimmer. 2023. Simplistic collection and labeling practices limit the utility of benchmark datasets for Twitter bot detection. In Proceedings of the ACM web conference 2023 . 3660–3669

  27. [35]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778

  28. [36]

    Olivier J Hénaff, Robbe LT Goris, and Eero P Simoncelli. 2019. Perceptual straightening of natural videos.Nature neuroscience 22, 6 (2019), 984–991

  29. [37]

    Carlos Javier Hernandez-Castro and Arturo Ribagorda. 2010. Pitfalls in CAPTCHA design and implementation: The Math CAPTCHA, a case study. computers & security 29, 1 (2010), 141–157

  30. [38]

    Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. 2022. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303 (2022)

  31. [39]

    Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022. Video diffusion models. Advances in Neural Information Processing Systems 35 (2022), 8633–8646

  32. [40]

    Yaosi Hu, Chong Luo, and Zhenzhong Chen. 2022. Make it move: controllable image-to-video generation with text descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18219–18228. 20 Lehao Lin, Ke Wang, Maha Abdallah, and Wei Cai

  33. [41]

    Guodong Huang, Chuan Ma, Ming Ding, Yuwen Qian, Chunpeng Ge, Liming Fang, and Zhe Liu. 2023. Efficient and low overhead website fingerprinting attacks and defenses based on TCP/IP traffic. In Proceedings of the ACM Web Conference 2023 . 1991–1999

  34. [42]

    Montree Imsamai and Suphakant Phimoltares. 2010. 3D CAPTCHA: A next generation of the CAPTCHA. In 2010 International Conference on Information Science and Applications. IEEE, 1–8

  35. [43]

    Iat Long Iong, Xiao Liu, Yuxuan Chen, Hanyu Lai, Shuntian Yao, Pengbo Shen, Hao Yu, Yuxiao Dong, and Jie Tang. 2024. OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Ling...

  36. [44]

    Rutvij H Jhaveri, Sankita J Patel, and Devesh C Jinwala. 2012. DoS attacks in mobile ad hoc networks: A survey. In 2012 second international conference on advanced computing & communication technologies . IEEE, 535–541

  37. [45]

    Junfeng Jing, Shenjuan Liu, Gang Wang, Weichuan Zhang, and Changming Sun. 2022. Recent advances on image edge detection: A comprehensive review. Neurocomputing 503 (2022), 259–271

  38. [46]

    Hongwen Kang, Kuansan Wang, David Soukal, Fritz Behr, and Zijian Zheng. 2010. Large-scale bot detection for search engines. In Proceedings of the 19th international conference on World wide web . 501–510

  39. [47]

    Mohammad Karami, Youngsam Park, and Damon McCoy. 2016. Stress testing the booters: Understanding and undermining the business of DDoS services. In Proceedings of the 25th International Conference on World Wide Web . 1033–1043

  40. [48]

    Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi

  41. [49]

    Suzi Kim and Sunghee Choi. 2019. Dotcha: A 3d text-based scatter-type captcha. In Web Engineering: 19th International Conference, ICWE 2019, Daejeon, South Korea, June 11–14, 2019, Proceedings 19 . Springer, 238–252

  42. [50]

    Kurt Alfred Kluever and Richard Zanibbi. 2009. Balancing usability and security in a video CAPTCHA. In Proceedings of the 5th Symposium on Usable Privacy and Security . 1–11

  43. [51]

    Karel Kubicek, Jakob Merane, Ahmed Bouhoula, and David Basin. 2024. Automating Website Registration for Studying GDPR Compliance. In Proceedings of the ACM on Web Conference 2024 . 1295–1306

  44. [52]

    Pierre Laperdrix, Gildas Avoine, Benoit Baudry, and Nick Nikiforakis. 2019. Morellian analysis for browsers: Making web authentication stronger with canvas fingerprinting. In Detection of Intrusions and Malware, and Vulnerability Assessment: 16th International Conference, DIMV...

  45. [53]

    Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. 2021. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 7331–7341

  46. [54]

    Jiangtao Li, Ninghui Li, XiaoFeng Wang, and Ting Yu. 2009. Denial of service attacks and defenses in decentralized trust management.International Journal of Information Security 8 (2009), 89–101

  47. [55]

    Linjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Ce Liu, and Lijuan Wang. 2023. Lavender: Unifying video-language understanding as masked language modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 23119–23129

  48. [56]

    Qiujie Li. 2015. A computer vision attack on the ARTiFACIAL CAPTCHA. Multimedia Tools and Applications 74 (2015), 4583–4597

  49. [57]

    Xigao Li, Babak Amin Azad, Amir Rahmati, and Nick Nikiforakis. 2023. Scan Me If You Can: Understanding and Detecting Unwanted Vulnerability Scanning. In Proceedings of the ACM Web Conference 2023 . 2284–2294

  50. [58]

    Xigao Li, Babak Amin Azad, Amir Rahmati, and Nick Nikiforakis. 2021. Good bot, bad bot: Characterizing automated browsing activity. In 2021 IEEE symposium on security and privacy (sp) . IEEE, 1589–1605

  51. [59]

    Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, et al. 2024. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177 (2024)

  52. [60]

    Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3431–3440

  53. [61]

    Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. 2023. Videofusion: Decomposed diffusion models for high-quality video generation. arXiv preprint arXiv:2303.08320 (2023)

  54. [62]

    Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2024. CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation. In Findings of the Association for Computational Linguistics ACL 2024 . 9097–9110

  55. [63]

    Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424 (2023)

  56. [64]

    Raman Maini and Himanshu Aggarwal. 2009. Study and comparison of various image edge detection techniques. International journal of image processing (IJIP) 3, 1 (2009), 1–11

  57. [65]

    Peter Matthews, Andrew Mantel, and Cliff C Zou. 2010. Scene tagging: image-based CAPTCHA using image composition and object relationships. In Proceedings of the 5th ACM Symposium on Information, Computer and Communications Security . 345–350

  58. [66]

    Manar Mohamed, Niharika Sachdeva, Michael Georgescu, Song Gao, Nitesh Saxena, Chengcui Zhang, Ponnurangam Kumaraguru, Paul C Van Oorschot, and Wei-Bang Chen. 2014. A three-way investigation of a game-captcha: automated attacks, relay attacks and usability. In Proceedings of th...

  59. [67]

    Vu Duc Nguyen, Yang-Wai Chow, and Willy Susilo. 2012. Breaking a 3D-based CAPTCHA scheme. In Information Security and Cryptology-ICISC 2011: 14th International Conference, Seoul, Korea, November 30-December 2, 2011. Revised Selected Papers 14 . Springer, 391–405

  60. [68]

    Vu Duc Nguyen, Yang-Wai Chow, and Willy Susilo. 2014. On the security of text-based 3D CAPTCHAs. Computers & security 45 (2014), 84–99

  61. [69]

    Nick Nikiforakis, Alexandros Kapravelos, Wouter Joosen, Christopher Kruegel, Frank Piessens, and Giovanni Vigna. 2013. Cookieless monster: Exploring the ecosystem of web-based device fingerprinting. In 2013 IEEE Symposium on Security and Privacy . IEEE, 541–555

  62. [70]

    Behzad Ousat, Esteban Schafir, Duc C Hoang, Mohammad Ali Tofighi, Cuong V Nguyen, Sajjad Arshad, Selcuk Uluagac, and Amin Kharraz

  63. [71]

    Nitisha Payal, Nidhi Chaudhary, and Parma Nand Astya. 2012. JigCAPTCHA: An Advanced Image-Based CAPTCHA Integrated with Jigsaw Piece Puzzle using AJAX. International Journal of Soft Computing and Engineering (IJSCE) 2, 5 (2012), 2231–2307

  64. [72]

    M Kameswara Rao, MSVK Maniraj, and B Sneha Ganga. 2014. Improved video captcha. Journal of Emerging Technologies in Web Intelligence 6, 4 (2014), 416–416

  65. [73]

    In Proceedings of the ACM on Web Conference 2024

    The Matter of Captchas: An Analysis of a Brittle Security Feature on the Modern Web. In Proceedings of the ACM on Web Conference 2024 . 1835–1846

  66. [74]

    Steven A Ross, J Alex Halderman, and Adam Finkelstein. 2010. Sketcha: a captcha based on line drawings of 3d models. In Proceedings of the 19th international conference on World wide web . 821–830

  67. [75]

    Philip Sedgwick. 2012. Pearson’s correlation coefficient. Bmj 345 (2012)

  68. [76]

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence 39, 6 (2016), 1137–1149

  69. [77]

    Dongyu She and Kun Xu. 2022. An image-to-video model for real-time video enhancement. In Proceedings of the 30th ACM International Conference on Multimedia. 1837–1846

  70. [78]

    Suphannee Sivakorn, Iasonas Polakis, and Angelos D Keromytis. 2016. I am robot:(deep) learning to break semantic image captchas. In 2016 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, 388–403

  71. [79]

    Asuman Senol, Alisha Ukani, Dylan Cutler, and Igor Bilogrevic. 2024. The Double Edged Sword: Identifying Authentication Pages and their Fingerprinting Behavior. In Proceedings of the ACM on Web Conference 2024 . 1690–1701

  72. [80]

    Oleg Starostenko, Claudia Cruz-Perez, Fernando Uceda-Ponga, and Vicente Alarcon-Aquino. 2015. Breaking text-based CAPTCHAs with variable word and character orientation. Pattern Recognition 48, 4 (2015), 1101–1112

  73. [81]

    Wenhao Sun, Rong-Cheng Tu, Jingyi Liao, and Dacheng Tao. 2024. Diffusion Model-Based Video Editing: A Survey.arXiv preprint arXiv:2407.07111 (2024)

  74. [82]

    Ray Smith. 2007. An overview of the Tesseract OCR engine. In Ninth international conference on document analysis and recognition (ICDAR 2007) , Vol. 2. IEEE, 629–633

  75. [83]

    Mengyun Tang, Haichang Gao, Yang Zhang, Yi Liu, Ping Zhang, and Ping Wang. 2018. Research on deep learning techniques in breaking text-based captchas and designing image-based captcha. IEEE Transactions on Information Forensics and Security 13, 10 (2018), 2522–2537

  76. [84]

    Isakwisa Gaddy Tende, Kentaro Aburada, Hisaaki Yamaba, Tetsuro Katayama, and Naonobu Okazaki. 2021. Development and Evaluation of Swahili Text Based CAPTCHA. In 2021 IEEE 3rd Global Conference on Life Sciences and Technologies (LifeTech) . IEEE, 293–297

  77. [85]

    Mayumi Takaya, Yusuke Tsuruta, and Akihiro Yamamura. 2013. Reverse Turing Test using Touchscreens and CAPTCHA. J. Wirel. Mob. Networks Ubiquitous Comput. Dependable Appl. 4, 3 (2013), 41–57

  78. [86]

    Upthrust. 2024. Runway, Luma, Kling, Pika, and Haiper: AI Video Generators Review Roundup. (August 2024). https://upthrust.co/2024/08/runway- luma-kling-pika-and-haiper-ai-video-generators-review-roundup Accessed: 2024-10-02

  79. [87]

    Luis Von Ahn, Benjamin Maurer, Colin McMillen, David Abraham, and Manuel Blum. 2008. recaptcha: Human-based character recognition via web security measures. Science 321, 5895 (2008), 1465–1468

  80. [88]

    Sheng Tian and Tao Xiong. 2020. A generic solver combining unsupervised learning and representation learning for breaking text-based captchas. In Proceedings of The Web Conference 2020 . 860–871

  81. [89]

    Jiawei Wang, Liping Yuan, and Yuchen Zhang. 2024. Tarsier: Recipes for Training and Evaluating Large Video Description Models. arXiv preprint arXiv:2407.00634 (2024)

  82. [90]

    Ke Wang, Lehao Lin, Maha Abdallah, and Wei Cai. 2025. Where is the Boundary? Understanding How People Recognize and Evaluate Generative AI-extended Videos. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems

  83. [91]

    Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Kevin Qinghong Lin, Satoshi Tsutsui, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, et al. 2023. All in one: Exploring unified video-language pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...

  84. [92]

    Tingting Wang and Jørgen Bøegh. 2014. Multi-layer CAPTCHA based on Chinese character deformation. In Trustworthy Computing and Services: International Conference, ISCTCS 2013, Beijing, China, November 2013, Revised Selected Papers . Springer, 205–211

  85. [93]

    Simon S Woo, Jingul Kim, Duoduo Yu, and Beomjun Kim. 2017. Exploration of 3D texture and projection for new CAPTCHA design. InInformation Security Applications: 17th International Workshop, WISA 2016, Jeju Island, Korea, August 25-27, 2016, Revised Selected Papers 17 . Springe...

  86. [94]

    Ping Wang, Haichang Gao, Xiaoyan Guo, Chenxuan Xiao, Fuqi Qi, and Zheng Yan. 2023. An experimental investigation of text-based captcha attacks and their robustness. Comput. Surveys 55, 9 (2023), 1–38

  87. [95]

    Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, and Yu-Gang Jiang. 2023. A survey on video diffusion models. arXiv preprint arXiv:2310.10647 (2023)

  88. [96]

    Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. Videoclip: Contrastive pre-training for zero-shot video-text understanding. arXiv preprint arXiv:2109.14084 (2021)

  89. [97]

    Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua. 2024. Can i trust your answer? visually grounded video question answering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13204–13214. 22 Lehao Lin, Ke Wang, Maha Abdallah, and Wei Cai

  90. [98]

    Xin Xu, Lei Liu, and Bo Li. 2020. A survey of CAPTCHA technologies to distinguish between human and computer. Neurocomputing 408 (2020), 292–307

  91. [99]

    Takumi Yamamoto, J Doug Tygar, and Masakatsu Nishigaki. 2010. CAPTCHA using strangeness in machine translation. In 2010 24th IEEE International Conference on Advanced Information Networking and Applications . IEEE, 430–437

  92. [100]

    Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. 2024. PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning. arXiv:2404.16994 [cs.CV] https://arxiv.org/abs/2404.16994

  93. [101]

    Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023. A survey on multimodal large language models. arXiv preprint arXiv:2306.13549 (2023)

  94. [102]

    Junnan Yu, Xuna Ma, and Ting Han. 2017. Usability investigation on the localization of text captchas: take chinese characters as a case study. In Transdisciplinary Engineering: A Paradigm Shift . IOS Press, 233–242

  95. [103]

    Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2024. MiniCPM-V: A GPT-4V Level MLLM on Your Phone. arXiv preprint arXiv:2408.01800 (2024)

  96. [104]

    Jerrold H Zar. 2005. Spearman rank correlation. Encyclopedia of biostatistics 7 (2005)

  97. [105]

    Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021. Merlot: Multimodal neural script knowledge models. Advances in neural information processing systems 34 (2021), 23634–23651

  98. [106]

    Zhengqing Yuan, Yixin Liu, Yihan Cao, Weixiang Sun, Haolong Jia, Ruoxi Chen, Zhaoxu Li, Bin Lin, Li Yuan, Lifang He, et al. 2024. Mora: Enabling generalist video generation via a multi-agent framework. arXiv preprint arXiv:2403.13248 (2024)

  99. [107]

    Shijie Zhang and Jong-Hyouk Lee. 2019. Double-spending with a sybil attack in the bitcoin decentralized network. IEEE transactions on Industrial Informatics 15, 10 (2019), 5715–5722

  100. [108]

    Ziyi Zhang, Shuofei Zhu, Jaron Mink, Aiping Xiong, Linhai Song, and Gang Wang. 2022. Beyond bot detection: combating fraudulent online survey takers. In Proceedings of the ACM Web Conference 2022 . 699–709

  101. [109]

    Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  102. [112]

    Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022. Learning to prompt for vision-language models. International Journal of Computer Vision 130, 9 (2022), 2337–2348

  103. [2023]

    In Proceedings of the IEEE/CVF International Conference on Computer Vision

    Text2video-zero: Text-to-image diffusion models are zero-shot video generators. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 15954–15964

  104. [2024]

    In Proceedings of the ACM on Web Conference 2024

    Medusa: Unveil Memory Exhaustion DoS Vulnerabilities in Protocol Implementations. In Proceedings of the ACM on Web Conference 2024 . 1668–1679

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.