Pith. sign in

Paper Citation Record · LEDGER

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

As of 15 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.11968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11968 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:09:24.884191Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ff6a5f7-d099-43b7-993d-ac873a99dab8 · outbound

This paper cites Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reference-guided ver- dict: Llms-as-judges in automatic evaluation of free-form text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.019458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.019458Z digest=sha256:5f8c054d4883cff01d78b51b92ee1b06b66b0cb815c77f41df0764d52bb4afe3

Observation e16c073f-8f7d-44a4-b9f4-20b48050c655 · outbound

This paper cites Qwen Technical Report.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.108626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.108626Z digest=sha256:884aeb144cd598aef4f50f14b22bf7b4605bf2a8cfa31ef3cce8654ff95c0ba9

Observation 8ed9183b-c458-4eff-943c-aeab1b085286 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Qwen2.5-vl technical report, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.210391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.210391Z digest=sha256:8601b6cf76a2e0bad0e72acb9a446e932544a8eedde5f2c10c632b0c68d51195

Observation d28aeb4f-a308-4d4c-aa2b-d3965368af78 · outbound

This paper cites Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mind the trojan horse: Image prompt adapter enabling scalable and decep- tive jailbreaking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.439134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:20.360131Z digest=sha256:c027ab596e1e956644d77506bcf32727f2918d56337e5496622d6c950b070369

Observation 66c0c188-e221-428c-89c6-419e70695c42 · outbound

This paper cites Cross-modal causal relation alignment for video question grounding.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal causal relation alignment for video question grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.311137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:20.453573Z digest=sha256:64012347235e32eb5d8c04b995efe9bea6193f2a4b55e323f763ee826224c157

Observation 291eba03-f8ca-4e32-8383-290a1468cd14 · outbound

This paper cites `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:20.679512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:20.679512Z digest=sha256:3c1c4ebeafdefce1672341824ba239bf83498fb58b061822a696f1aed56eabf3

Observation ddb8e526-1796-4dcb-b0bc-df01372b2d6c · outbound

This paper cites Automated hate speech detection and the prob- lem of offensive language.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Automated hate speech detection and the prob- lem of offensive language

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:31.155231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:20.802326Z digest=sha256:d6c213e5468c5dfa297e81e4d19bf40b391fa874b86f7d6238c4d9515249612a

Observation b73647cf-4c09-453c-a52e-b06da2025cfb · outbound

This paper cites Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Adversaflow: Visual red team- ing for large language models with multi-level adversarial flow

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.943692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:20.986593Z digest=sha256:d9b219b85d90896ae3d3bc05bd84a733f2ef90e103720a19ba2453c75ddd2846

Observation 5dd2071b-9d0f-4494-a9da-7fb14055249b · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.129340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.129340Z digest=sha256:4602c4ac1b379ccecaf2bd529fe8df70e31e0f10959e3b3f15682a805f72362e

Observation 1d0e4511-6370-4ab8-b552-78b526a0fc98 · outbound

This paper cites Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Figstep: Jailbreaking large vision-language models via typo- graphic visual prompts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.762839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:21.237058Z digest=sha256:66cbd4f17ce1f5202e3737944559b94d56637c1f5277b30b0c9cbe795f67fd9a

Observation 03018278-1c7c-4639-adce-5a6329ee0fcc · outbound

This paper cites Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Agent smith: a single image can jailbreak one million multimodal llm agents exponentially fast

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.596842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:21.381323Z digest=sha256:70474daf9152e43538b877931aeb7ea14e0cdf25fa3ed9265c95812d75f45a86

Observation 3226d307-bc2f-41c6-8b47-466275d0936c · outbound

This paper cites Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:26.210420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:21.503096Z digest=sha256:b33d589878cb815d2c2b5a0c6436723604ab7dafe41a62be58ee4bf12ace6fcd

Observation 85cf70de-8463-44ab-83dd-6938be8f6a6d · outbound

This paper cites Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.645814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.645814Z digest=sha256:04b0a9c328fdd9d7bcd6ae4f9622c0a481804fc7b5a8349b53a7a776f00827cc

Observation f7ff1dbd-f1a1-4405-ac42-2ef80dee1729 · outbound

This paper cites Curiosity-driven Red-teaming for Large Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Curiosity-driven Red-teaming for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.745442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.745442Z digest=sha256:8101f5ab0674df43fcffc33d5c61d50d4ab3e6abfcbea7b9f6c415105fa05d2d

Observation 3799aeeb-ca05-4bfd-aff4-9d4345b8f798 · outbound

This paper cites Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Videojail: Exploiting video-modality vulnerabilities for jail- break attacks on multimodal large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.321298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:21.813983Z digest=sha256:d6b46be583fcde19f6175f73f7d09a8882d5b865a438cd4f5290156173a061d7

Observation f3453251-37a2-4b02-bc7c-43a2825fee6f · outbound

This paper cites GPT-4o System Card.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:21.874258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:21.874258Z digest=sha256:75694aa7d99d298ba6880034b2871a0d55583d11ca9cd879c47e8750180984fd

Observation 143932b7-7b54-4fbe-a676-f6a160177086 · outbound

This paper cites Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:30.122424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:21.954102Z digest=sha256:f457222fc02934aa0c45b6b4a7ccfab4fb1cc50c621f564dbf5edf38aa5b5ed4

Observation 77f3739a-8d88-483a-adf3-2f34ef7daae7 · outbound

This paper cites an unresolved cited work.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:09:29.886185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:22.063009Z digest=sha256:ef7f0d4c2bcf86e83c92a004de068afaa81b795c4a663415122b17f0e21a27e8

Observation ab245714-1a27-4a3d-95bc-c4c9376f8cd0 · outbound

This paper cites Mixtral of Experts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Mixtral of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.151616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.151616Z digest=sha256:87b7cb8fedf8e1806a1e7a88afad5de39267048822401128b7574b552b54217a

Observation 23bfdaf8-b932-4062-b746-684a6ff46bce · outbound

This paper cites Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.262209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.262209Z digest=sha256:d478b43b1383e58bb769d60cb052ffc22143fff246912b72710ce51149dd7334

Observation c4fd354b-7b66-47ce-8fce-1aa561f06e9f · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.633249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:22.481992Z digest=sha256:8252cf5e6f2c2c37b03003e1fb6d93bbe5849e72345085c7370bee0791891027

Observation 26fcdc41-bd95-4bb1-b0d9-aef4fda1bea2 · outbound

This paper cites Learning diverse attacks on large language models for robust red-teaming and safety tuning.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.585003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.585003Z digest=sha256:32484ecd7976389982538500826753e64e048fa817e4e0597d0ad0bcf09db3b8

Observation 59d4164c-800a-4cb4-80c1-86a6a3160165 · outbound

This paper cites FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:25.713243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:22.681672Z digest=sha256:709a529cfa48d7182e35c9b9320d7a80b7b6d24a4d8f78708cd708105979a08a

Observation f6c3f2bb-6f94-4c4c-9ce0-9d04a072b29f · outbound

This paper cites Red Teaming Visual Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Red Teaming Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.747756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.747756Z digest=sha256:5d3ad6879743ed01d25ef264be60e1b60a424e276f9ac4fd19594b5955009107

Observation 33ef3c8a-7a3e-4d2a-bca0-a74912568c55 · outbound

This paper cites Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Images are achilles’ heel of alignment: Exploit- ing visual vulnerabilities for jailbreaking multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.412453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:22.819473Z digest=sha256:8118c46c175e5e5e2bf691e3acf4b4e89ca44b3501b2b3d44661962e768ccbc8

Observation d0329f0b-e3b1-4d21-85ae-fa4993207658 · outbound

This paper cites GroundingGPT: Language en- hanced multi-modal grounding model.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GroundingGPT: Language en- hanced multi-modal grounding model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:29.191945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:22.889336Z digest=sha256:336ab7e87f6c8fa464a52ae45b1c0926a3e467a804660b202a72bf6199cb0717

Observation c7154d59-2136-4e09-a0b6-26e205f80d7e · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.005461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.005461Z digest=sha256:cb989e0481f1577d785c2265bef8a6c29a208b66cc17f800d603ab8718db618e

Observation d9571e64-afe9-4fd3-a968-abaea2c7befa · outbound

This paper cites Visual instruction tuning, 2023.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Visual instruction tuning, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.904324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:23.166951Z digest=sha256:9ee2303d95cca9265ff37ae8e538806ca06fe064b9142d52503942f4ba3828cd

Observation ea1be7ab-bb2f-4614-a63e-0968bef7222a · outbound

This paper cites Prompt Injection attack against LLM-integrated Applications.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Prompt Injection attack against LLM-integrated Applications

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.315417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.315417Z digest=sha256:f95d2f1066cb1e9dc56859df26bbb886c39c6d382a5b7a8812e4c2ac8219833b

Observation 018af706-99bd-4699-bd4c-474e63673e0f · outbound

This paper cites Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.643297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:23.404182Z digest=sha256:9bf4c8442e23d6aadf883dddd5c6d9dd39ccd73b558ce24bd0b00c21d3a19659

Observation b6e18b6e-9118-4450-bd3c-723107c9b59d · outbound

This paper cites LLaMA 4: Advancing Multimodal Intelli- gence.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA 4: Advancing Multimodal Intelli- gence

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.459147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:23.520377Z digest=sha256:52850962bd21907b19e96f5c2cb967c09e012a7320a0336d3719a718dddc0378

Observation 531a2f32-3929-4b2e-9217-77d748f5e1b7 · outbound

This paper cites Jailbreaking Attack against Multimodal Large Language Model.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Jailbreaking Attack against Multimodal Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:23.717332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:23.717332Z digest=sha256:8ebe2e3274454b0d7b2ffc74e88f5a0f545e27e311adbae0f860724ba6538781

Observation fcdc0549-e95f-4248-831d-5d17c73f418d · outbound

This paper cites Gpt-4o system card, 2024.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gpt-4o system card, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:28.219937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:23.877523Z digest=sha256:cb8ca2b56437458b58cfd857daf003622d5e8a82bf1ca68b24f8e826095fd1f8

Observation 6b7e1439-2684-4bc1-9960-364cf76dbbd1 · outbound

This paper cites Cross-modal attention congruence regularization for vision-language relation align- ment.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Cross-modal attention congruence regularization for vision-language relation align- ment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.964083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:23.970611Z digest=sha256:9b958ba1c5ea0d06d1a00a17d34539244615d3dafbaab542dffd7aa0c8381a23

Observation fac8730f-908b-48c4-93c5-9c6ed93363b0 · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.061030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.061030Z digest=sha256:7e1b19cfe089b0d898ba33fcec497b2ff1e8631ef6f7f94e9038af46be5d613d

Observation 5539a05b-bf7b-4e23-96a9-bcf4dc4462f3 · outbound

This paper cites On the adversarial robustness of multi-modal foundation models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation On the adversarial robustness of multi-modal foundation models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.713195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:24.135215Z digest=sha256:c323ff4068f0a43ad34f938a1463326f8948fe53fb9c3bbb8876d7fd1bd38b71

Observation 19ad089f-5c2d-4bb0-a906-1e0c1954c34e · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemini: A family of highly capable multi- modal models, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.507691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:24.223768Z digest=sha256:6539ce0d3a0b64b6bb056d7dbb7977a82128737aaaefd46bf0f3e5693e647836

Observation 738711b2-8710-4033-9fec-60cc76de8881 · outbound

This paper cites Gemma 3 technical report, 2025.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Gemma 3 technical report, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.281820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:24.302682Z digest=sha256:1b775497f65578372e208aac3835a7836ae60f381b5c6043a6203c76c7c8f8f6

Observation ce20cddb-4b6f-4e99-9e13-e06abfd1f244 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.373479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.373479Z digest=sha256:e6c3b298b81a462ef74ceef316ccb2718b3388609d02fc9b2e58c3a5ca7d91ea

Observation 807d9dd6-23a9-47e5-8aa9-fa257b9aedaf · outbound

This paper cites Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Stop reasoning! when multimodal llm with chain-of-thought rea- soning meets adversarial image, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:27.092656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:24.470625Z digest=sha256:18af9b1d8dbb80113e016b5095a2cc3c42efad432d43072ae6d5f893e336c7dd

Observation 8c0cf09f-e4bb-4809-83e5-b7bffefece4f · outbound

This paper cites Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:09:25.245696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:24.580550Z digest=sha256:48bc3fd32a7b315ea2299ad730b4af1f802658832b9f6c3fdca615f72a3ca48f

Observation 07a803cf-9713-49e5-ae02-241870041023 · outbound

This paper cites Distraction is all you need for multimodal large language model jailbreaking.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Distraction is all you need for multimodal large language model jailbreaking

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:09:26.829247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:09:24.672129Z digest=sha256:2e90b0570230df5146644ad25e23d456fe6ad715047be86c833ccc396004d350

Observation 8bdb0029-ddc9-4f19-9bd5-1ce547d883c3 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.745035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.745035Z digest=sha256:9840717db01dae19ed12d59e9216539c3154e3688c209e9900ea794d436d041f

Observation 23f244aa-ebae-4ab8-ac21-295c9395e1d8 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.806832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.806832Z digest=sha256:38a10dc3fb6234480b3b4de2b1e24dacf516479c650a844e81a860ad8fc8331e

Observation 1612cb63-f7b3-4493-b43e-c44ffd16bb53 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.884191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.884191Z digest=sha256:343c2443796dd555a87c8f88537dca3ad7a7304473c31ccc707f03b3153fe80b

Pith citing papers

No inbound Pith citation observations are available.