Pith. sign in

Paper Citation Record · LEDGER

SafeVid: Toward Safety Aligned Video Large Multimodal Models

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2505.11926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11926 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:49:49.493020Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:13:57.233182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1ad4736-3e7f-4374-a09c-536b091ccff7 · outbound

This paper cites URLhttps://huggingface.co/datasets/ gretelai/gretel-safety-alignment-en-v1.

SafeVid: Toward Safety Aligned Video Large Multimodal Models URLhttps://huggingface.co/datasets/ gretelai/gretel-safety-alignment-en-v1

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.221604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.269141Z digest=sha256:d92050663e700155c1dc1e566729d1c69327c34340b58b16366c7ee8ef4faff9

Observation 1f554298-3b44-49df-b871-86107684589d · outbound

This paper cites GPT-4 Technical Report.

SafeVid: Toward Safety Aligned Video Large Multimodal Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.274039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.274039Z digest=sha256:2624741d7c17772bec3f2f188cf2434473847f7dd49731387de896578133218d

Observation f4c3b572-a298-4802-bba4-74a67503792d · outbound

This paper cites Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.279644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.279644Z digest=sha256:207c4500c0f77b40644678cef891f843bbd96e299004f8cf571a0dc5d72595ae

Observation 9755fe92-a1bb-42a2-9b16-895775525571 · outbound

This paper cites Claude.https://claude.ai/chats, 2023.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Claude.https://claude.ai/chats, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.208950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.283989Z digest=sha256:d81b1a19a03e074734162dc9804d61e41e80b725b7167ce0de63f9c3b4eca06b

Observation 49337657-b6bf-46b4-90f3-079beae8302c · outbound

This paper cites Qwen Technical Report.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.287871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.287871Z digest=sha256:cce957ef07b613ab0010d8b23ac46d91852412aae6c3c16dedc6b2ae75be63e8

Observation 749056c8-5454-4745-9001-dd7237aa4e44 · outbound

This paper cites Qwen2.5-VL Technical Report.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.292153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.292153Z digest=sha256:54219ee9604e49e93b7462b880e090c2d6860fa6330feacc124d017bade60ef8

Observation 8dfb6d81-d0e4-4f58-bb83-8a5c4234aa3e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Constitutional AI: Harmlessness from AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.297385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.297385Z digest=sha256:afa508569b33501f4ee8fbc2220ea5e2c8880c75f59cc733bb4aa0a939eeb968

Observation bb599b88-6d53-4036-89bc-c720a7a329c1 · outbound

This paper cites Safeinfer: Context adaptive decoding time safety alignment for large language models.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Safeinfer: Context adaptive decoding time safety alignment for large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.196776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.302204Z digest=sha256:956b10220e520255fed859a9a7b3f6afdac0d6b2ebcd5f1a3958cb843d24993a

Observation d07f0f92-9076-483a-841c-843fb3299a0f · outbound

This paper cites Movieclip: Visual scene recognition in movies.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Movieclip: Visual scene recognition in movies

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.184265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.306036Z digest=sha256:2c513bc3d8489eec76c6b5fb7560776d3d875e39b16c947b92e7c4b14d758559

Observation 1c27fa4e-6254-458c-aa08-798f766380a6 · outbound

This paper cites Internlm2 technical report, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Internlm2 technical report, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.171461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.309796Z digest=sha256:dcdc30863c79414e3dba0b84faa215dc864db798d286dc8b68903754fa94324c

Observation 1bc8e709-dbcf-4ae7-8743-9fdb5d3ec250 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.313768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.313768Z digest=sha256:4de9f478f341277606ca17bff5a4386eb0d8a06b8c69eb6c320d11b2fc334ae1

Observation efc923aa-0142-4905-a9fa-930b2d2320b3 · outbound

This paper cites Bypassing Safety Guardrails in LLMs Using Humor.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Bypassing Safety Guardrails in LLMs Using Humor

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:49:49.801421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.317848Z digest=sha256:be6336888e5bf477675fd81fb266cdfdc8b200abb30dc60fea0898f8f4f2149f

Observation 3eaa028f-c90a-4dee-a94c-7f725ce779e9 · outbound

This paper cites Large scale holistic video understanding.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Large scale holistic video understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.152016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.321787Z digest=sha256:1ec1f040e1797e10221a73109c0754d14d29c165aaad0975bbafb57c240c92d3

Observation 5dc728eb-2b3f-4b94-93cf-c0df3b668c1f · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.140513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.325314Z digest=sha256:a2aba4157294d39e1d7247293ad1320543d1583bd664a374caf36d5787cf08d7

Observation bef74d9b-e54e-4377-8408-7d7f50aff464 · outbound

This paper cites Videojail: Exploiting video-modality vulnerabilities for jailbreak attacks on multimodal large language models.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Videojail: Exploiting video-modality vulnerabilities for jailbreak attacks on multimodal large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.127468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.328916Z digest=sha256:a59e28b0ce140c9133e5be4403b90bf730d9f337012f7635d3252d64713b3cf9

Observation 43d4ddd0-d980-47dc-9568-d6dca537a604 · outbound

This paper cites Flames: Benchmarking value alignment of llms in chinese.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Flames: Benchmarking value alignment of llms in chinese

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.114284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.332568Z digest=sha256:9f9c636d82ca19f1949f3181ce1750d46b47f2c40b8ff5d782c4c6fe62df7a9f

Observation e9b536b1-e2df-4a50-88c4-7b0484e426fd · outbound

This paper cites GPT-4o System Card.

SafeVid: Toward Safety Aligned Video Large Multimodal Models GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.336287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.336287Z digest=sha256:c61171331fc567d84ef157ae6b59ff6e62cf24e223061410222dc4f10268737b

Observation 4a63f49f-1614-4236-ae60-8c3b9bd954d8 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.NeurIPS, 2023.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Beavertails: Towards improved safety alignment of llm via a human-preference dataset.NeurIPS, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.101310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.340701Z digest=sha256:01a93b43fe0ae79a6d4edec6fdf4aad502196bcedd041e97b04808d18796da21

Observation f9cf6536-2396-4a68-94db-4790b442b590 · outbound

This paper cites Pku-saferlhf: A safety alignment preference dataset for llama family models.arXiv e-prints, pages arXiv–2406, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Pku-saferlhf: A safety alignment preference dataset for llama family models.arXiv e-prints, pages arXiv–2406, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.088917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.344664Z digest=sha256:80d82fa3a7346913519d76ec32c3be79f8b552816bd4c83396b4271c79639401

Observation 6fd37571-ad30-40d3-beb4-445be51ea9a3 · outbound

This paper cites Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.348937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.348937Z digest=sha256:4676c66eff6810f8e6268f47410d9361514eafd90bafcbba1c0611341adb56e3

Observation 3bd1b770-d493-4bd2-aea7-ec3638b774ac · outbound

This paper cites Awesome-llm-robotics, 2022.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Awesome-llm-robotics, 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.076356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.352946Z digest=sha256:c26be8a545577a73a3f2bb5cf1620dc799d7c92764a655cd7e4e5aa57b8f7fd4

Observation c6503423-5272-4aa4-b6d0-0fad007be127 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

SafeVid: Toward Safety Aligned Video Large Multimodal Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.357305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.357305Z digest=sha256:a8958325d55eafa59cfd014983204edc323044cc658f10e4706137e2f6ca3501

Observation 87996247-0660-48ae-a16f-955e0e181d22 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SafeVid: Toward Safety Aligned Video Large Multimodal Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.361201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.361201Z digest=sha256:a3e6926db0c961d7a99c92b19e0dff8cdb662f58edd4905284f5892dabaf6cea

Observation cd5c1096-4d41-4f38-b657-6a2bec385181 · outbound

This paper cites DeepSeek-V3 Technical Report.

SafeVid: Toward Safety Aligned Video Large Multimodal Models DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.365357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.365357Z digest=sha256:1f29f405dcc1ea0a9ce47a69618dde34550f358e8e201fc8899a88b16b47f7e4

Observation ecdaefe5-8393-4bf3-bdeb-8a5dd5c773c4 · outbound

This paper cites Harnessing llms for automated video content analysis: An exploratory workflow of short videos on depression.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Harnessing llms for automated video content analysis: An exploratory workflow of short videos on depression

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.063049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.369122Z digest=sha256:ed452c4d605c687cd2c415b86c65ceaaefeba5f8f3eeb4932a70fe8a9fd1e8fd

Observation a267a1d3-36b9-4788-b4fc-f06d4ccf50e3 · outbound

This paper cites Mm-safetybench: A benchmark for safety evaluation of multimodal large language models.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Mm-safetybench: A benchmark for safety evaluation of multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.049668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.373600Z digest=sha256:4f8b22f05df66464a466aa21f2881c9e8e4233f733ad8a799f5e1493e4ffa38b

Observation 774099fb-cbb7-4933-8ea2-de38ef1c1508 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.377610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.377610Z digest=sha256:c6ed8be6d82967ccc01a2f01d98149f1ae19286c623de5e898850e0184be0e46

Observation 7f1ef541-78d0-4e95-91a8-c1c737eec1e1 · outbound

This paper cites Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks.arXiv e-prints, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks.arXiv e-prints, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.035464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.381416Z digest=sha256:777496e3b5a29dcf03dd9b4cbdb317d8e16c9f23a4f161d0a560c0d614bd4e44

Observation 84f96ebd-3586-462b-a13b-4756df2cd196 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.385121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.385121Z digest=sha256:42cb8bf3200b3da460ac05c38ceccf87aad91e352c0e86b29693b74b4a5e83cc

Observation 7ddbf46d-e6f2-4f9d-b908-97f5830096ac · outbound

This paper cites Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types.NeurIPS, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types.NeurIPS, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.021549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.389161Z digest=sha256:787743d22593524d61e4b2209f292243156aa161fb2dc360d00b2b6e046372eb

Observation a6782bb9-ce57-437e-9b32-02360622c919 · outbound

This paper cites Chatgpt.https://chat.openai.com/chat, 2023.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Chatgpt.https://chat.openai.com/chat, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:50.000483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.393240Z digest=sha256:0d524f953bec2d0c6f3167072f880b3650806fa075441b2a177b4ab07a780057

Observation 2cbdd0da-ed37-4065-845c-49897e409b8d · outbound

This paper cites Training language models to follow instructions with human feedback.NeurIPS, 2022.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Training language models to follow instructions with human feedback.NeurIPS, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.396866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.396866Z digest=sha256:f363d21f5f6087d1204c977f61727ec92c6532d623a4496ba37c592bcae047ce

Observation 10c40284-b4bf-4048-a8b9-44aa5adc2ca1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.400709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.400709Z digest=sha256:47d1002d854c395ed8603fb60de10a46b98c170cc1d430e5f98d96fc23522b62

Observation c11ab2f9-4fdb-4e4d-bddd-2d24434d55a4 · outbound

This paper cites Safetywashing: Do ai safety benchmarks actually measure safety progress?NeurIPS, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Safetywashing: Do ai safety benchmarks actually measure safety progress?NeurIPS, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:49.972098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.404651Z digest=sha256:64d415c640f4c09c6e87cc6fa2226e96bea00c65c5c2e9b9cc7b9f570b5676a9

Observation 93e2d6c3-663b-4afb-9cf8-0f6fa4de5584 · outbound

This paper cites Large Language Model Safety: A Holistic Survey.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Large Language Model Safety: A Holistic Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.408428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.408428Z digest=sha256:e874bccb8aec853e0f8c79e5bb59384e0b8607058c3ac38e0460ef9710083877

Observation e4dda6f1-ae08-4407-9b67-7c753f65d8fc · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

SafeVid: Toward Safety Aligned Video Large Multimodal Models A StrongREJECT for Empty Jailbreaks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.412207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.412207Z digest=sha256:711b36a2df2e51b7bfa33abfbbcef64a3641a49027c3072b14b33656a62fce52

Observation 59e8a7ae-3d92-4030-9502-67e93709b04d · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.416402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.416402Z digest=sha256:c6bfb293c43854a38194fe1dbf9462859115890f34e678cf5d9c4a1ca2ac03d5

Observation 63e0bef6-deae-45b9-a20e-a419c2755135 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.420281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.420281Z digest=sha256:e579bcc6c2c5521af75e173c0e2f0235b88df4cffcd0f8b99eb5c0ca9c6293b6

Observation c893b13c-09e9-4cd4-aef5-2b8b5883464f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SafeVid: Toward Safety Aligned Video Large Multimodal Models LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.424246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.424246Z digest=sha256:d630bf1239b445b91ded23b4ace398ee7b4da3b50f475c80ff3f3530b285e215

Observation 2b1926a5-84e9-43b1-99be-ae980477d9a5 · outbound

This paper cites Ideator: Jailbreaking large vision-language models using themselves.arXiv preprint arXiv:2411.00827, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Ideator: Jailbreaking large vision-language models using themselves.arXiv preprint arXiv:2411.00827, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.428709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.428709Z digest=sha256:da2c9a9ebc3c92cc8e622bc18d14c690bf03995f687d07f915958aea8824d7ca

Observation 8a02c5f9-9ab5-47d4-b94f-cc3fd0cee864 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

SafeVid: Toward Safety Aligned Video Large Multimodal Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.432833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.432833Z digest=sha256:953b46ae20e1e8905ae3e837b30cf69febe7745af789e1f57a33e09e0b5ba226

Observation e2f56e1b-0682-4489-b391-d459133a36ea · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Internvideo2: Scaling foundation models for multimodal video understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.438126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.438126Z digest=sha256:9f3624a4c58620a29a409b00d3c82f1c32eb60d5328db2196b66f8d924dfbd45

Observation e35c34c0-d8e4-40ff-8075-557ac3875cd8 · outbound

This paper cites Fake alignment: Are llms really aligned well? InNAACL, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Fake alignment: Are llms really aligned well? InNAACL, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:49.952659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.442438Z digest=sha256:c566b728929e02906e24cc3bf347f4e8412426c2e672511b147d210c56b79f9d

Observation 3ea6302c-aa86-4567-a215-32da17f40578 · outbound

This paper cites Gpt4video: A unified multimodal large language model for lnstruction-followed understanding and safety-aware generation.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Gpt4video: A unified multimodal large language model for lnstruction-followed understanding and safety-aware generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:49.940315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.447429Z digest=sha256:f85d86f23f332bd385b2f1f68f5500cec7536ee987ce1d3d7ed5845fa2315f53

Observation 60d04b8e-f8ff-4970-8432-cee4a9b6cd40 · outbound

This paper cites Jailbroken: How does llm safety training fail?NeurIPS, 2023.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Jailbroken: How does llm safety training fail?NeurIPS, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:49.928450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.451919Z digest=sha256:daa4d0f750881f38cbd2978cd5326f0f1e14138bacf109549ffd8fcd5b538917

Observation fc8c26be-4ac3-4ed9-aadb-16e92869d2de · outbound

This paper cites Videoclip: Contrastive pre-training for zero- shot video-text understanding.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Videoclip: Contrastive pre-training for zero- shot video-text understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:49.915901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.455771Z digest=sha256:02cd1a6097d592de88b1cd871bbeed0464eef3da03552fc505aa49dce7c3b544

Observation 9d985729-cf7a-4a2c-8b8b-249824dc2b58 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SafeVid: Toward Safety Aligned Video Large Multimodal Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.459361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.459361Z digest=sha256:62f0e823d0ba343d8e8d1bc65bc1d8c0ce4cf62a603630f76826bfe65f5f66fd

Observation 6831ce43-0974-4404-bd89-b022ae622c13 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.464120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.464120Z digest=sha256:5a68abf87767b322f7dfc43d37a0517977e17fb2d42ad79b8139af250450e20b

Observation ac69b33e-7093-439a-a44a-7a0269c42e65 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Llava-next: A strong zero-shot video understanding model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.469109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.469109Z digest=sha256:ed18b5677a46fa8e8a91c36709fc8e1841c10bb7d3324b2c8cf5aa3cb12fa080

Observation cdd644ca-8c10-464f-b2dd-67bc22e8e367 · outbound

This paper cites Spa-vl: A comprehensive safety preference alignment dataset for vision language model, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Spa-vl: A comprehensive safety preference alignment dataset for vision language model, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:49.895543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.473099Z digest=sha256:7547b55086a14e14a81823eb0b1f29cf7210d39391d82f0da4ae3b3c43d982a1

Observation 9c183fa0-ef71-4989-83ac-ef253deb6fec · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, 2024.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Llava-next: A strong zero-shot video understanding model, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:49.883593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.477432Z digest=sha256:7421e874a7604ac8b4362057cc2fbd45170ca6cd04e133811b5e54f67f2e43d0

Observation fb1984f9-60b1-449a-bd55-baba4736575a · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:49:49.870515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:49:49.481284Z digest=sha256:c87447c9673cf55f6dc9393d966dd29abc1b250e1c03b4b80f55750839cccf0b

Observation 1b497a68-7bff-4ef0-8468-78430bcd844d · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Fine-Tuning Language Models from Human Preferences

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.489432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.489432Z digest=sha256:01bf2adfa853c295e95737f7061cc62729c242cd2a15b820c383c0309504e3cb

Observation 1711eeb6-f979-466d-9a7d-8d254fcc51b5 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SafeVid: Toward Safety Aligned Video Large Multimodal Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.493020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.493020Z digest=sha256:7d9858c8a2413fe26d96dcbb96d47846f063ba92b14f30285452ec2b0c218cec

Observation f9288a81-0878-488e-b574-c7f39e59fd2a · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

SafeVid: Toward Safety Aligned Video Large Multimodal Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:49.485430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:49.485430Z digest=sha256:6fd2c3ae28da0a5aa5d77168e62b9cbf0f30eb721566e867b4db2bee2156537d

Pith citing papers

Observation 9c61bd1e-6b1d-4570-b708-f7f5e7f71cec · inbound

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework cites this paper.

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework SafeVid: Toward Safety Aligned Video Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T13:13:57.233182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:13:57.233182Z digest=sha256:0572a32f70cf548e6e767fcb0259fcc45761bfc29477006711718a7bb516f824