Pith. sign in

Paper Citation Record · LEDGER

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision

As of 12 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2412.12722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12722 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:52:06.558105Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:18:51.206793Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T08:17:36.629263Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f91bcbe3-94bf-4516-891a-ec631693f6e8 · outbound

This paper cites GPT-4 Technical Report.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.399038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.399038Z digest=sha256:6c100eb3853e1f3fe94cc95eb2677195c6296b95ab44097609098071c6941b1b

Observation 07c7b448-4eb4-46ff-82c0-d8656569421f · outbound

This paper cites Qwen2.5-VL Technical Report.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.412507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.412507Z digest=sha256:43e5150cc02045f8eebf33b5350e6a6f74c0c0fa594680bc914f93d31f4384de

Observation a7a6eb8c-ef97-4fdd-b457-acac2c0c6d05 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.418197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.418197Z digest=sha256:b8ba244f891f8abbd596f4065ffe2b40679d712ecffd9e4e6e1a2b91cd9db24a

Observation d6a13326-9ff3-46f8-b0eb-72686175bc09 · outbound

This paper cites And in Table 10, we demonstrate the statistical significance using a t-test, with p-values consistently less than 0.05, confirming their significance.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision And in Table 10, we demonstrate the statistical significance using a t-test, with p-values consistently less than 0.05, confirming their significance

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:52:07.267093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:52:06.558105Z digest=sha256:6f14f9ff71dd5484aa22d6e27619a29679b61d10a3a4aec47fadb9c16e653079

Observation 977aa6ab-e8cc-4dc0-a45c-a83c7871c729 · outbound

This paper cites VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.435588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.435588Z digest=sha256:09a36dedaa9032c44e3a4b376f5e4b75bd029ad7323d5a6c9598f11c2da264e4

Observation 40fc9fdb-9720-4fb6-8076-796de0e50df8 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.441334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.441334Z digest=sha256:21262fa18da43e3ff84d19d87a05d67eb770b4aa23609a6c65884e1090e6fcef

Observation b9ef07a7-d525-4177-8901-260be3fe3e69 · outbound

This paper cites MirrorCheck: Efficient Adversarial Defense for Vision-Language Models.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision MirrorCheck: Efficient Adversarial Defense for Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.446381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.446381Z digest=sha256:da5775d3689a1ad776720278feab4a985b142d64474f8032587b0d828c262245

Observation 7b1736b7-ac9a-4c0e-8d82-b797a7d691ff · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.451631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.451631Z digest=sha256:57da875fe43c058dcd6ab331d1530fb4267ee1ffd1845389e9b63e9c9d8a72cc

Observation 1ff9fd07-66ba-4f93-92af-82769d679bc3 · outbound

This paper cites Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.457092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.457092Z digest=sha256:147323dd88bc26e35be3b694430e10c72aaf74d5c7ba73d6febbc2834bea59fd

Observation ecc7c7e8-cf79-4bbd-8d72-16db1904e85c · outbound

This paper cites an unresolved cited work.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.462433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.462433Z digest=sha256:2d6bb3b189355589397578fcb501b5a0b4db00bacdc23669bc3d6799bd644431

Observation 8e6ddbc5-15e5-47e2-8f9d-9e220b70407d · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.467601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.467601Z digest=sha256:d0e4dcc1524d160e86011de8fa3a7a7e91fd5a0a8a1aa88d236f2abc606cb1db

Observation 79fd509e-e8af-4387-993b-e99923945b17 · outbound

This paper cites Crafting papers on machine learning.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Crafting papers on machine learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.472783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.472783Z digest=sha256:317d77e8d9f8a14da3d700b6b846fadcc421eb227da9b8f0e098294f284998fc

Observation 9c907a8f-6ff2-4370-8cdf-1f9011ab0d71 · outbound

This paper cites Red Teaming Visual Language Models.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Red Teaming Visual Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.478022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.478022Z digest=sha256:0957e17b0b7050cf5a847373fc49cb8fb319e528fc63822f59d63c4edde0167e

Observation ca673e37-1847-4c5e-81fc-948085644f94 · outbound

This paper cites Interpreting and Mitigating Hallucination in MLLMs through Multi-agent Debate.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Interpreting and Mitigating Hallucination in MLLMs through Multi-agent Debate

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.484604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.484604Z digest=sha256:823bcd7cf80914446e43820e99404062b4ce30c5a7bb5dd28e815028cd407a0f

Observation 64a1b9dd-4f1f-4667-bedd-474ae84b0984 · outbound

This paper cites Compromising Embodied Agents with Contextual Backdoor Attacks.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Compromising Embodied Agents with Contextual Backdoor Attacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.489864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.489864Z digest=sha256:f76392eda15a89c01b1f15eac3f08b6df9c4fa3b20db5d01b793a1574e51a678

Observation 193def2a-7b71-4835-8ff9-56d4142b2fed · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.505659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.505659Z digest=sha256:4a7a0da8a604dcdbe0e5526ad08a4861dbcf0b6f336107c9c2ddf8da8894d9e2

Observation d8cd0a61-a02c-4aad-a3d9-4a2b65ee099a · outbound

This paper cites Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.510672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.510672Z digest=sha256:0182fbd17a7538147ead669a7d8340ebc17aa912df5a43ca10141d6a63d3d9f8

Observation f19e412a-8dab-4f56-a416-a9e83a48042e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.515912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.515912Z digest=sha256:7435c45da5e6a5991af66a8c2e74f32fe1d29076dcb973e2132521c1fb715064

Observation 2434febd-cfce-489f-9b6b-f4e12f6a35e3 · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.521313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.521313Z digest=sha256:631bda383e3539f0da9656331ebe2e40398f55f1f3a298f817a880a9462fe79c

Observation fd6c90b7-9b6e-4aaf-8fc4-9b34e96372f1 · outbound

This paper cites A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.526634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.526634Z digest=sha256:406c4d7772a67f50c7e6591cdc94d1d3a21352ff0406d837c4b18d2416708c27

Observation 0ca53795-b073-4c7c-9897-dff41b740ee3 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.537012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.537012Z digest=sha256:9323e863332f44eee80e88756ff464de56b675c8f05264e1aa281a5fa8644460

Observation ab662ad0-9046-4ac6-8eff-add245f9eadc · outbound

This paper cites JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.542812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.542812Z digest=sha256:eb375f7b6b88bf935fd15ae8f2948d3bd7d892b424da41f68590103ea79cb653

Observation 47e8d337-8f74-4057-acbd-26d4bfc8cd67 · outbound

This paper cites RTA-100 is a real-world typographic attack dataset, in which the handwritten tag from incorrect classes is placed next to the objects in the image.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision RTA-100 is a real-world typographic attack dataset, in which the handwritten tag from incorrect classes is placed next to the objects in the image

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:52:07.308907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:52:06.547824Z digest=sha256:1496ecef5cc837a32cf5023a7457795e73141db6421d1174735820f9883bb720

Observation 315ab618-47e5-4b99-84fb-d3bf2bbc54e1 · outbound

This paper cites Both MM-safetyBench and HADES are datasets for evaluating LVLM in safety-critical scenarios.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Both MM-safetyBench and HADES are datasets for evaluating LVLM in safety-critical scenarios

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:52:07.286936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:52:06.552746Z digest=sha256:598d173e60873ac71005803b447cf213075e64a61dd2eb8e5507ca4b4fc844e1

Observation 30a71ed2-5819-4d5e-9412-d3d826c96962 · outbound

This paper cites MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.500384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.500384Z digest=sha256:b1ae7c70da7efa1eb4e4d2979da9330e9738f19a24c28e3996ec16a9f84d20dc

Observation 3469c160-a225-4e6a-bbdb-0b86803b98b3 · outbound

This paper cites Feature squeezing: Detecting adversarial examples in deep neural networks.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Feature squeezing: Detecting adversarial examples in deep neural networks

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:52:07.332672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T13:52:06.531864Z digest=sha256:283ff236f7747ac9222a728df697a59d530bbfc0c243740614c9fcc901e93e26

Observation dfdd65c6-6f08-4d00-8d8d-5c21a5ee1b48 · outbound

This paper cites M., Vedaldi, A., Zisserman, A., and Jawahar, C.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision M., Vedaldi, A., Zisserman, A., and Jawahar, C

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.495063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.495063Z digest=sha256:76070efbbccb479f32762b984963a4775bdb0234d20182e4bcde7124a5abd39f

Observation 65fc6f2a-e77c-46cf-a783-5238f208294b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.405908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.405908Z digest=sha256:fd8403e70ebf8e15ae7eb297a03f2bbbf072b63899b789530eb3e2c9e00eed3a

Observation 49825eb2-d8e9-4cde-86e1-cd5cf62956d0 · outbound

This paper cites Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.429481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.429481Z digest=sha256:a00656873a8656ba6f36d7eadb64613290997380675f24e214cd93927470faed

Observation f6e69485-f1b8-4bfb-a5b2-c017f58bba82 · outbound

This paper cites SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.423820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.423820Z digest=sha256:6f2a5e3635ecd65542dc141f6a41056a0b5073c05e048622a2021f2830ec1087

Pith citing papers

Observation 6789da18-c782-4e69-b559-861ca7543d87 · inbound

A Survey on Training-free Alignment of Large Language Models cites this paper.

A Survey on Training-free Alignment of Large Language Models Defending LVLMs Against Vision Attacks through Partial-Perception Supervision

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:51.206793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:18:51.206793Z digest=sha256:5d451b6cb2d1751b00b647248f250d2750793585567dd46e1128e14f633ec337

Observation e0b10156-c474-44e0-817a-960849da6d09 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Defending LVLMs Against Vision Attacks through Partial-Perception Supervision

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.631161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:1bb1359ed124d1ad8b61283090be0f4ae8e9d97def23febe08084b0de97689a7