Pith. sign in

Paper Citation Record · LEDGER

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models

As of 21 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2509.00373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00373 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:44:56.708066Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T19:22:43.082083Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T19:23:40.845848Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0708b9c-3417-4db4-a5a7-af4755228e2e · outbound

This paper cites Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.560366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.560366Z digest=sha256:8405dc74d05b505f2012fff368de686c597c7bde444eae4f7825e1454aaf7700

Observation 3227f213-a1aa-4f19-81e2-84ec332bb600 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.686334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.686334Z digest=sha256:ed4bea1c947da5374bfaa37da4e9b054e9b05a50a2b5112d47fe59839fe3b53e

Observation abffdf20-d6b4-49db-b979-225b93a0e9c2 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.736911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.736911Z digest=sha256:d8275cdd486faa8bc1ae630e25872c2321d7e4b4c39c994edb73ccd40d7cad7c

Observation 348f2d69-e9be-4c67-a1e5-92c1f97f128c · outbound

This paper cites org/abs/2502.01042.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models org/abs/2502.01042

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.911424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.911424Z digest=sha256:4bf4a0a3f5dcb8c4ac977ade5bdeef66941f888eef75b82328e1116acaf4b901

Observation a4ea9379-7816-4e1a-ae76-8f8b17bb589c · outbound

This paper cites Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.970593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.970593Z digest=sha256:988566547a715f41e72afc6f75fde1856d5578c989ba97896372aa6b7e6f8ef8

Observation 4a534b28-1ab2-4c06-b50e-4a969a466e5e · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.020846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.020846Z digest=sha256:8499aa2f9ef1174f0034331d275d04997ca5ab7477e1ebf0e0ddd89ea04e26aa

Observation 95ee2fb0-ff21-425e-be10-c99c8bf9b066 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.091583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.091583Z digest=sha256:efa56fd38fecc975b02340aef99b871299262e1c4770c031270b3751d49a8771

Observation 3b940c23-c3cf-4361-8f61-dc2aa2003cf5 · outbound

This paper cites Training language models to follow instructions with human feedback.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Training language models to follow instructions with human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.147734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.147734Z digest=sha256:6c4f9f3e01b13166c19935f130b4c7092385c661c162fa4db83fc5afe91c81a0

Observation 91a1b601-a548-4dff-be26-deed374c2d07 · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.199824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.199824Z digest=sha256:849cd5c766f5f7b689bba20177879d79e28ed297e8b9afe94d5fdeb46ff04970

Observation 2a5973cb-a02f-4be9-a475-36a9752362cd · outbound

This paper cites Visual Adversarial Examples Jailbreak Aligned Large Language Models.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Visual Adversarial Examples Jailbreak Aligned Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.261344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.261344Z digest=sha256:4f4772f48f5a4b533cc4ae20d48f388f112cc3c69cf37c8a744015f9c3ff3167

Observation f7401d2b-0a18-4308-9b08-2dd5c8c7f73a · outbound

This paper cites On the Adversarial Robustness of Multi-Modal Foundation Models.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models On the Adversarial Robustness of Multi-Modal Foundation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.319661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.319661Z digest=sha256:18c6e7c3e0d7e8622eb95087f969993ed66fac6f821e1ccc8c9fd171986eb18d

Observation ddd9d865-c49b-4625-bf38-8cd50f50e119 · outbound

This paper cites Learning to summarize from human feedback.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Learning to summarize from human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.450480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.450480Z digest=sha256:4a32a02f9d52a05020f4987123f9209fa37e0e877e3dbd1627e0c32edbbd231a

Observation 646136b6-51c0-4fa4-95bb-153dfd9e1ad3 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models LaMDA: Language Models for Dialog Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.510535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.510535Z digest=sha256:196ca811f282d0b5b604903621632e2abaddf320d05367230ae40533039f30ea

Observation 0c001538-34a8-42f6-9972-7f97dea0a2fb · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Fine-Tuning Language Models from Human Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.600389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.600389Z digest=sha256:0df3407d7487e21c31f1ed07432b39de03d31ab20214560fc6ed7efab8c3801d

Observation a5373c59-f4d4-4435-8a6e-e0ef15286754 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.661410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.661410Z digest=sha256:965db7832b58752834511e769e226a6a2ff767eaf830f396fa449c7f003dde62

Observation 5fbaf813-d142-41e0-8e51-156b1b7b5221 · outbound

This paper cites Understanding and Rectifying Safety Perception Distortion in VLMs.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Understanding and Rectifying Safety Perception Distortion in VLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.708066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.708066Z digest=sha256:6d7e387e05958ddde3b96ddd2f0c51f4c338a774832e2e9a875106d87bb1d864

Observation 5fedac94-aaaf-46b5-9e1d-666e7eed10a6 · outbound

This paper cites semanticscholar.org/CorpusID:125209808.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models semanticscholar.org/CorpusID:125209808

Reference 1952

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:44:57.927865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T13:44:55.505326Z digest=sha256:13137cced6b17766e9db5263207dd872fa4d9dc0d2fb190963bc809f9a5894d2

Observation 7705a784-fd13-4e0a-bbfd-754c2a849208 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:56.384261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:56.384261Z digest=sha256:2662ce2c41f9901c54d3c214aabc26a5129c94d43da57671a36c68ce0f5cbc72

Observation 11528131-98cc-44e4-9a0f-d85f6e415ec6 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.794893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.794893Z digest=sha256:d4cf3d8bbbcbbf119a9b55cacb1d63537f8577e3b9177c6a09399fc90e926ac4

Observation 6fe0083e-acf1-4012-91ef-9c0ad18e8301 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models A General Language Assistant as a Laboratory for Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.370516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.370516Z digest=sha256:09608eb97c94909eb8dd28fd52cca3e2f49eaba9a3793ea19066570a1d7b4864

Observation 997c3dc8-d86b-472d-abb7-23ae05a9a131 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.439302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.439302Z digest=sha256:725b1adcc3f270e388ebe21e3c557170e86d72129be1692b0e2631d0c8d37d2a

Observation 4fd7a290-1e3c-4ead-8c4d-74ec6625fbee · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.630606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.630606Z digest=sha256:45237be12712b8846ef82bcde4209a1380bb830a759d1e68bb7939c5e414783a

Observation 6cc7a34f-f324-447c-82fc-6cdea44c5e2d · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models Refusal in Language Models Is Mediated by a Single Direction

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.307039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.307039Z digest=sha256:a144e825a5b74fddf2cc8fc0070bfa5cb22c09adb43cc446f94a95f6f7afb15e

Observation c4b7143d-a5e2-4c76-abb8-656d7930f615 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:55.849334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:55.849334Z digest=sha256:f541ee464a44ab383aab4bbcf072743cf9ae1f9c2386a579e4234c807523754e

Pith citing papers

Observation f4736f01-8815-4b13-9e3e-564cff6477db · inbound

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models cites this paper.

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:23:40.847621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T19:22:43.082083Z digest=sha256:616d9a3d0c2abadd2edb5106de364e7897ae4bf5c5666b679cdbac17fe5162f9