Pith. sign in

Paper Citation Record · LEDGER

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.11968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11968 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:09:24.884191Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ff6a5f7-d099-43b7-993d-ac873a99dab8 · outbound

This paper cites Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.019458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.019458Z digest=sha256:c52de9b55dcf98ff2524d786a788fcb9fa79c68ef65ec4058b4ab430659ecb98

Observation e16c073f-8f7d-44a4-b9f4-20b48050c655 · outbound

This paper cites Qwen Technical Report.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.108626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.108626Z digest=sha256:9c19dfb2cb4c6ebefcdd73ca87d192133fb2019a7ba0bfba9326f4f7e46c7844

Observation 8ed9183b-c458-4eff-943c-aeab1b085286 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen2.5-vl technical report, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.210391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.210391Z digest=sha256:d2f5917990c2a0d05a97aaf364778db49746ed9df14cee4c0d641e8b9157b28b

Observation d28aeb4f-a308-4d4c-aa2b-d3965368af78 · outbound

This paper cites Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.439134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:20.360131Z digest=sha256:8ebdad93fdd9237e1f12ab29f33977b759c406809ef7e460f75204b5b202f1e5

Observation 66c0c188-e221-428c-89c6-419e70695c42 · outbound

This paper cites Cross-modal causal relation alignment for video question grounding.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal causal relation alignment for video question grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.311137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:20.453573Z digest=sha256:2cc98b8c2fe2b223082ae8fa914453ab0251c45c96984de2b0338abd842a8e6a

Observation 291eba03-f8ca-4e32-8383-290a1468cd14 · outbound

This paper cites `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.679512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.679512Z digest=sha256:b34413c7a64baa781f2618966eabca8943ff7644fbbb45d0aae625458dce30cc

Observation ddb8e526-1796-4dcb-b0bc-df01372b2d6c · outbound

This paper cites Automated hate speech detection and the prob- lem of offensive language.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Automated hate speech detection and the prob- lem of offensive language

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.155231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:20.802326Z digest=sha256:34225936d61d5177276900769f8d8e6cf9b27daed416debd648de21a46ecdece

Observation b73647cf-4c09-453c-a52e-b06da2025cfb · outbound

This paper cites Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.943692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:20.986593Z digest=sha256:bd1f107d9e700a5a083b3f0017ce21c36403e766d068f2853cb1bbaf98e52bab

Observation 5dd2071b-9d0f-4494-a9da-7fb14055249b · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.129340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.129340Z digest=sha256:f6516fad108744315190a237037bd406f316a6c569b6ac0cd57e98ac0091def0

Observation 1d0e4511-6370-4ab8-b552-78b526a0fc98 · outbound

This paper cites Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.762839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.237058Z digest=sha256:ff1f4d03d3c8aa8c4e467e972d2cf30f0ab1dc33d6ab08a74df1641b739ea3a2

Observation 03018278-1c7c-4639-adce-5a6329ee0fcc · outbound

This paper cites Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.596842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.381323Z digest=sha256:1be1103bd8812086ebc786c9a4dca3962c807f133fa57638bc868304ad3f4e23

Observation 3226d307-bc2f-41c6-8b47-466275d0936c · outbound

This paper cites Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:26.210420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.503096Z digest=sha256:0492d47c6d0d710bc3949bf8b273e4db948f3f4c41d0a7d7037048f2df10cc26

Observation 85cf70de-8463-44ab-83dd-6938be8f6a6d · outbound

This paper cites Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.645814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.645814Z digest=sha256:c14031479b1a5baffdddea63ca8807951fdd1b87a15691e775f82b8602af51af

Observation f7ff1dbd-f1a1-4405-ac42-2ef80dee1729 · outbound

This paper cites Curiosity-driven Red-teaming for Large Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Curiosity-driven Red-teaming for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.745442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.745442Z digest=sha256:48eab9192d7012526cb62ab30dddbf33d7240e4bae21e624f6d76bc7ca01bb63

Observation 3799aeeb-ca05-4bfd-aff4-9d4345b8f798 · outbound

This paper cites Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.321298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.813983Z digest=sha256:3632b179ddcd7238ca804d13dfcd992091584e0ec8d52ab8ae10aeb40cf56a5e

Observation f3453251-37a2-4b02-bc7c-43a2825fee6f · outbound

This paper cites GPT-4o System Card.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.874258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.874258Z digest=sha256:db95e2bb21a17c36ec858b1958a60a961742af6b2f1d69e70606038ef4ef9898

Observation 143932b7-7b54-4fbe-a676-f6a160177086 · outbound

This paper cites Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.122424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:21.954102Z digest=sha256:a45dfa356b27ba39605e2f95c68a6a36df4cdef9d40801c7053ff7e0a3736c94

Observation 77f3739a-8d88-483a-adf3-2f34ef7daae7 · outbound

This paper cites an unresolved cited work.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:09:29.886185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.063009Z digest=sha256:0fa2f9315e258eae64a766343e9b2ffc8d4d9450ba762091cb20c2b6f7f33e30

Observation ab245714-1a27-4a3d-95bc-c4c9376f8cd0 · outbound

This paper cites Mixtral of Experts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mixtral of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.151616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.151616Z digest=sha256:de9c648047bf2f3e7f5040ee1af7f9cea0c764d0b8fe3295131a7ab7da433c92

Observation 23bfdaf8-b932-4062-b746-684a6ff46bce · outbound

This paper cites Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.262209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.262209Z digest=sha256:425e9c8fc4c809ac258228d08403f621d2fa0620e5041425ea68333858c25d8b

Observation c4fd354b-7b66-47ce-8fce-1aa561f06e9f · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.633249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.481992Z digest=sha256:3ada4ac41eb029d30ad302a2445362f6e5b23c935abfa60eeedf791e7b611807

Observation 26fcdc41-bd95-4bb1-b0d9-aef4fda1bea2 · outbound

This paper cites Learning diverse attacks on large language models for robust red-teaming and safety tuning.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.585003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.585003Z digest=sha256:15491500a60d5973f9f4b64095b938547d156bb900087b7425c9dbe8eab406be

Observation 59d4164c-800a-4cb4-80c1-86a6a3160165 · outbound

This paper cites FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:25.713243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.681672Z digest=sha256:98be8b79b8015e2bfbef6ab19183142b7295744241588d21b4c630f700736c7b

Observation f6c3f2bb-6f94-4c4c-9ce0-9d04a072b29f · outbound

This paper cites Red Teaming Visual Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Red Teaming Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.747756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.747756Z digest=sha256:cc693c7cb798f29b12ab711780f8a9edee678eae0b4394f638ed228e6f80a80d

Observation 33ef3c8a-7a3e-4d2a-bca0-a74912568c55 · outbound

This paper cites Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.412453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.819473Z digest=sha256:de0e9e00a417eb1bcefd916817f4a165f4686be420aba500ee62b28da81e4b4e

Observation d0329f0b-e3b1-4d21-85ae-fa4993207658 · outbound

This paper cites GroundingGPT: Language en- hanced multi-modal grounding model.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GroundingGPT: Language en- hanced multi-modal grounding model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.191945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:22.889336Z digest=sha256:cca4dae3dc9bc0e5393a7d8613a1921d834811c8b1508e5ce5bfa3609642fc19

Observation c7154d59-2136-4e09-a0b6-26e205f80d7e · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.005461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.005461Z digest=sha256:b69a131c5d5236ace8675248ddf7d795642ae1af45ec2643e56a6782fcf63ced

Observation d9571e64-afe9-4fd3-a968-abaea2c7befa · outbound

This paper cites Visual instruction tuning, 2023.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Visual instruction tuning, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.904324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.166951Z digest=sha256:3ffc8d76a592e335cd0b3dc75e60c23b14fb374a046f4535667249c882eaf27e

Observation ea1be7ab-bb2f-4614-a63e-0968bef7222a · outbound

This paper cites Prompt Injection attack against LLM-integrated Applications.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Prompt Injection attack against LLM-integrated Applications

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.315417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.315417Z digest=sha256:165343c7037a1814852d9a40557dfacbeac731b0c27f5c533f2590a6d21f4f04

Observation 018af706-99bd-4699-bd4c-474e63673e0f · outbound

This paper cites Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.643297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.404182Z digest=sha256:98d02a96a8b016b86a4fc1f7693c035d99dbac45ef7e22c4a9a45ca202d51ea3

Observation b6e18b6e-9118-4450-bd3c-723107c9b59d · outbound

This paper cites LLaMA 4: Advancing Multimodal Intelli- gence.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA 4: Advancing Multimodal Intelli- gence

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.459147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.520377Z digest=sha256:317f84ae1d7f42913160763d775856dadc392c7df9069592670d97833a3109bc

Observation 531a2f32-3929-4b2e-9217-77d748f5e1b7 · outbound

This paper cites Jailbreaking Attack against Multimodal Large Language Model.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Jailbreaking Attack against Multimodal Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.717332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.717332Z digest=sha256:dab72422840150a841527b0d7f095219e5c168c55c786df0206d1a8a3a47d2de

Observation fcdc0549-e95f-4248-831d-5d17c73f418d · outbound

This paper cites Gpt-4o system card, 2024.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gpt-4o system card, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.219937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.877523Z digest=sha256:030eb613dbe75af8d9d17771dd963a3a3c0598b1efa02f4324413695775555af

Observation 6b7e1439-2684-4bc1-9960-364cf76dbbd1 · outbound

This paper cites Cross-modal attention congruence regularization for vision-language relation align- ment.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal attention congruence regularization for vision-language relation align- ment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.964083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:23.970611Z digest=sha256:f83ccf9a2d1cd50c391e7454384d562a5d29bbd5d4959a7910dd646637afd73b

Observation fac8730f-908b-48c4-93c5-9c6ed93363b0 · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.061030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.061030Z digest=sha256:f9492b330a52874160701eebcdeb20a62210fc67feebeaafeda7205309a6d57b

Observation 5539a05b-bf7b-4e23-96a9-bcf4dc4462f3 · outbound

This paper cites On the adversarial robustness of multi-modal foundation models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation On the adversarial robustness of multi-modal foundation models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.713195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.135215Z digest=sha256:9de888dff81bec92f78e7132e94b1d45eb4db6344c5192f4cc236fcbac649edf

Observation 19ad089f-5c2d-4bb0-a906-1e0c1954c34e · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemini: A family of highly capable multi- modal models, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.507691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.223768Z digest=sha256:640112c810fe85473db7949f03ea54fa8e1db0abd0452856ea17dc97ebb96ebe

Observation 738711b2-8710-4033-9fec-60cc76de8881 · outbound

This paper cites Gemma 3 technical report, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemma 3 technical report, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.281820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.302682Z digest=sha256:bfe4642c6ba681e0ba374337e1822a8538211708c61aa8260884853c73d8c6ba

Observation ce20cddb-4b6f-4e99-9e13-e06abfd1f244 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.373479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.373479Z digest=sha256:de8a2792c25a8010273f5c4fe5375e46cd988d138247dd0df496bbd9b0c401e8

Observation 807d9dd6-23a9-47e5-8aa9-fa257b9aedaf · outbound

This paper cites Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.092656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.470625Z digest=sha256:9f9da88f9a9fb7fa90d37ceed9b20c3fc56a226a8bbc9d5289f973751b6314fc

Observation 8c0cf09f-e4bb-4809-83e5-b7bffefece4f · outbound

This paper cites Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:25.245696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.580550Z digest=sha256:e16c4c970fa3940fffee6d040ccc8114bd230c30044b264a9a428b80274910c9

Observation 07a803cf-9713-49e5-ae02-241870041023 · outbound

This paper cites Distraction is all you need for multimodal large language model jailbreaking.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Distraction is all you need for multimodal large language model jailbreaking

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:26.829247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:09:24.672129Z digest=sha256:5fa4250d253002b66de8de515ee816774542c66ea8d64369b32c9b3e281444e3

Observation 8bdb0029-ddc9-4f19-9bd5-1ce547d883c3 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.745035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.745035Z digest=sha256:95bb945a4a01723dfb813c0a4d0239c07da5de3b39320e7a389791a331e6505e

Observation 23f244aa-ebae-4ab8-ac21-295c9395e1d8 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.806832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.806832Z digest=sha256:46c6d32eb224251ed60dfef34cfae5464d4624b5b5b4412a1ce6b9ef98284776

Observation 1612cb63-f7b3-4493-b43e-c44ffd16bb53 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.884191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.884191Z digest=sha256:417f7162aeb5d4a7530a2647796655a9b14609aca38d069b8dbb65fb34d3c5a4

Pith citing papers

No inbound Pith citation observations are available.