Pith. sign in

Paper Citation Record · LEDGER

Multimodal Model Diffing for Feature Discovery and Control

As of 12 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2608.09928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09928 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:56.224644Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved74
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 63a68806-aeda-41aa-8974-f329e2ebba67 · outbound

This paper cites Pixtral 12b: A new frontier in image and text understanding.

Multimodal Model Diffing for Feature Discovery and Control Pixtral 12b: A new frontier in image and text understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.734741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.734741Z digest=sha256:c832e980a487fde7bed68a532066088a1ef4c02916d43fa91532e3234e76fc29

Observation e980897a-63ad-4ab4-9cae-ab1bace9e2e9 · outbound

This paper cites Golden gate Claude.

Multimodal Model Diffing for Feature Discovery and Control Golden gate Claude

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.740439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.740439Z digest=sha256:dbab1bf9e5c857e7b757722a334bc0a23d404c8f9d64fd8577a8f4485d124643

Observation b478f42b-d307-42c5-b16b-5954735aa5a1 · outbound

This paper cites SAE on activation differences.

Multimodal Model Diffing for Feature Discovery and Control SAE on activation differences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.744964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.744964Z digest=sha256:3ef6681abbf0b4ca85d6695020d4790df4dca8b1c5d075a03b63e9f1a13c8220

Observation eed6e3f5-9874-4c87-b480-11f0d9fdb4ad · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Multimodal Model Diffing for Feature Discovery and Control Refusal in Language Models Is Mediated by a Single Direction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.749442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.749442Z digest=sha256:7756e0ca38955e7da2b694a5f6787ecccfe6f81f9cc0b122d192fcbede1693bb

Observation 15f9a23f-fb05-4140-b13b-89fcd2bd4712 · outbound

This paper cites Revisiting model stitching to compare neural representations.Advances in neural information processing systems, 34:225–236, 2021.

Multimodal Model Diffing for Feature Discovery and Control Revisiting model stitching to compare neural representations.Advances in neural information processing systems, 34:225–236, 2021

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.754302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.754302Z digest=sha256:949b7f7b1ba7a612a51a6a2a5d02df8594700e5a135c96d34580eab18d03118c

Observation 99bad67a-8f11-4798-a8d3-1e4ee9868bcc · outbound

This paper cites Representation Topology Divergence: A Method for Comparing Neural Network Representations.

Multimodal Model Diffing for Feature Discovery and Control Representation Topology Divergence: A Method for Comparing Neural Network Representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.758852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.758852Z digest=sha256:2d8d7a10087ed1ed9926e27b10946ff15cea06bc9245463155ed2012567d08cc

Observation fb3300c3-6ac8-4870-bef9-0d7a1e819f96 · outbound

This paper cites Understanding information storage and transfer in multi-modal large language models.Advances in Neural Information Processing Systems, 37:7400–7426, 2024.

Multimodal Model Diffing for Feature Discovery and Control Understanding information storage and transfer in multi-modal large language models.Advances in Neural Information Processing Systems, 37:7400–7426, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.764254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.764254Z digest=sha256:49e38e8383917603e1c34367996d1cb88cd2e7e0e088af5666da9390a48b0b6a

Observation 775b35d7-5787-42c1-b28c-727175dec30e · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023.

Multimodal Model Diffing for Feature Discovery and Control Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.768697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.768697Z digest=sha256:fb1ec155faa7de90fbc180cc4d783d47bc06d960d72598e5f5148523b5d87630

Observation bfb2af8f-ac38-41fa-a848-228a761cca3a · outbound

This paper cites Stage-wise model diffing.

Multimodal Model Diffing for Feature Discovery and Control Stage-wise model diffing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.773486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.773486Z digest=sha256:3519f67c25bf2c3138af20e8309ddb4c3c75fa0bfdb8b4d23f2a27dfe0a4847e

Observation 96afae2d-fab5-435d-bf48-229175c2211e · outbound

This paper cites Observing and controlling features in vision-language-action models.arXiv preprint arXiv:2603.05487, 2026.

Multimodal Model Diffing for Feature Discovery and Control Observing and controlling features in vision-language-action models.arXiv preprint arXiv:2603.05487, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.778192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.778192Z digest=sha256:ba3875a8d8bce54c84da1e907ec4bd7e299555ea51fc272853a210d0bcafd4d7

Observation bae5706b-c4c3-4713-829f-3cfd8479df19 · outbound

This paper cites Improving Steering Vectors by Targeting Sparse Autoencoder Features.

Multimodal Model Diffing for Feature Discovery and Control Improving Steering Vectors by Targeting Sparse Autoencoder Features

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.782849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.782849Z digest=sha256:26758a74af5d8f0cc6d1a66f1959c785569f2a900887245738a90ba9618a8d09

Observation 21aef9b8-39d4-4a4a-bb5e-514a373a0a0d · outbound

This paper cites Pappas, Florian Tramer, Hamed Hassani, and Eric Wong.

Multimodal Model Diffing for Feature Discovery and Control Pappas, Florian Tramer, Hamed Hassani, and Eric Wong

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.788114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.788114Z digest=sha256:448251491b8ee208224736c6237639edd32350d87ff7a86d0c2b2f25c6fd7daf

Observation 943068ae-9db0-4d97-a492-eea47ed36f53 · outbound

This paper cites Interpreting and Controlling Vision Foundation Models via Text Explanations.

Multimodal Model Diffing for Feature Discovery and Control Interpreting and Controlling Vision Foundation Models via Text Explanations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.793237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.793237Z digest=sha256:a1d639d20cb0a7d515e4a9eb432d77f4d6302a69ba8b5a944b6c2bf1bf55daca

Observation 729e283c-fe58-406d-a371-03d8e94288e4 · outbound

This paper cites LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.798668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.798668Z digest=sha256:68a8607e39dd0cd86a874454ff3d2703be3e59782e7776ed7363544bf89e80b4

Observation 47c9ed4e-1705-4432-9bf6-b809e4061e02 · outbound

This paper cites Explaining How Visual, Textual and Multimodal Encoders Share Concepts.

Multimodal Model Diffing for Feature Discovery and Control Explaining How Visual, Textual and Multimodal Encoders Share Concepts

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:17:57.283192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:55.803988Z digest=sha256:9af007b21e6120d5664b6e56fea3f2bc708cb6ef54bd149248af835f18015c3c

Observation 9ea49434-9b1f-4ea9-b45e-0b22dc3bbf00 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Multimodal Model Diffing for Feature Discovery and Control Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.809262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.809262Z digest=sha256:e67f460546b3c45e0a09252b783bc22fdda5cc28c2a994296789f6dddfcb1b77

Observation ef347515-df97-4217-990c-bac9c5c5f962 · outbound

This paper cites Case study: Interpreting, manipulating, and controlling CLIP with sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Case study: Interpreting, manipulating, and controlling CLIP with sparse autoencoders

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.814349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.814349Z digest=sha256:f00bfb95cd807be91d960c3cde321c1cb326671db5b045161fc6b4cd562b01cd

Observation c03b7f1f-509a-432a-bf1a-565eab037242 · outbound

This paper cites Toy Models of Superposition.

Multimodal Model Diffing for Feature Discovery and Control Toy Models of Superposition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.819230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.819230Z digest=sha256:a401c9bf217e030e75a6052e8967ba1c1fdcaab3db1d64bb29407c320d69fdc3

Observation 907a2509-8501-4d4e-b1ce-4313ef05a681 · outbound

This paper cites Why does unsupervised pre-training help deep learning? 11:625–660, March.

Multimodal Model Diffing for Feature Discovery and Control Why does unsupervised pre-training help deep learning? 11:625–660, March

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.824530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.824530Z digest=sha256:c700087a5060cf246477999143e3303d29fcc958a20b62026122a38cb5e77dcb

Observation 5dc90ba2-6e4d-4bbf-98b4-ace19979d697 · outbound

This paper cites Interpreting CLIP's Image Representation via Text-Based Decomposition.

Multimodal Model Diffing for Feature Discovery and Control Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.829626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.829626Z digest=sha256:b669a44fa36acd15a057c7ca6bd4c356c41ac186d868ef835aa512e01e44e48b

Observation 4227ee35-1ef9-4683-9e9f-b03e81994704 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Scaling and evaluating sparse autoencoders

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.834643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.834643Z digest=sha256:9c13f1d25cf2a45946272caf3a4c60d217adce0077673a4df83f8a0b54cac583

Observation a67d0a03-cb2c-48bc-99b6-5077d2ca94a8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Multimodal Model Diffing for Feature Discovery and Control Gemma 2: Improving Open Language Models at a Practical Size

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.839825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.839825Z digest=sha256:3d80d4d42a2e5486b0dea671c6dc7277937846d63f5961602d623c339cc8eacb

Observation 2891bc5a-3300-4926-b5a2-2e673fc4e1d2 · outbound

This paper cites FigStep: Jailbreaking large vision-language models via typographic visual prompts.

Multimodal Model Diffing for Feature Discovery and Control FigStep: Jailbreaking large vision-language models via typographic visual prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.845475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.845475Z digest=sha256:d620363fbe57edf26cc78b2f6a7798210d241b1b26cd397f06e80f39d0477a65

Observation eb7628dc-8244-471c-97d8-776e74cf4700 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Multimodal Model Diffing for Feature Discovery and Control Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.850637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.850637Z digest=sha256:b5071391f6567a276d8c04c36f14f220f6604dbfa7f8731aef3f58a181cbfc29

Observation 383f8bd4-0f3d-4f1e-8663-f5c8552e8985 · outbound

This paper cites Not all features are created equal: A mechanistic study of vision-language-action models.

Multimodal Model Diffing for Feature Discovery and Control Not all features are created equal: A mechanistic study of vision-language-action models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.856252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.856252Z digest=sha256:1e678aa14c8186188cf5a834d6df4644d710a2013b9fe8d2e9419e567ad9ed77

Observation 594550ea-48d5-4b18-a375-73064b1a6c2e · outbound

This paper cites The Llama 3 Herd of Models.

Multimodal Model Diffing for Feature Discovery and Control The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.861229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.861229Z digest=sha256:e9ae2dcf7e7cbe86da86f8d3e2576308ba54831302a147368c9a19a1ba92d3e0

Observation 51b802d7-af4b-4cde-979e-97fc6d7c7e49 · outbound

This paper cites Mechanistic interpretability for steering vision-language-action models.

Multimodal Model Diffing for Feature Discovery and Control Mechanistic interpretability for steering vision-language-action models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.865869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.865869Z digest=sha256:e8a65fc0e1e874bcd85f49227fdb149b96f0a2b0aebe063d790bbd23f93fc778

Observation 3b640386-0590-4017-8b10-8dcab1ed9d2b · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.870806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.870806Z digest=sha256:45444ce651b9e9c61df75098aa8ddccd4ca953359db531de82e53428da01cba5

Observation e94666bd-7d15-4980-b0a4-2c789fe46bb3 · outbound

This paper cites In-context learning creates task vectors.

Multimodal Model Diffing for Feature Discovery and Control In-context learning creates task vectors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.875585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.875585Z digest=sha256:dacbbe893793af6d8b941453f9793b9b692fab6775c2b20f9061d9271b7f7959

Observation ed622c66-b2d4-4498-8016-67dc19488d7e · outbound

This paper cites VLSBench: Unveiling Visual Leakage in Multimodal Safety.

Multimodal Model Diffing for Feature Discovery and Control VLSBench: Unveiling Visual Leakage in Multimodal Safety

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.880064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.880064Z digest=sha256:8622597b6f4ac3bc25b728ef3f644818ed84e04c46d2781fb74ea80aabd53200

Observation 92e3a5d3-924f-4d22-961c-de23b2368db0 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Multimodal Model Diffing for Feature Discovery and Control Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.885168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.885168Z digest=sha256:523d316a5432032519765588a5940789af5e8cc39013f06b5aad7b4e293bc37e

Observation 60bfe481-7159-4238-b912-6accdf089da1 · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

Multimodal Model Diffing for Feature Discovery and Control Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.889674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.889674Z digest=sha256:7dfd3be43522ff8692a8860a3ac1c5fe809754f1832b7ce6cc9bb48c258dfa0c

Observation dd219166-8d27-475a-a1c4-06ce11c66684 · outbound

This paper cites A “diff” tool for AI: Finding behavioral differences in new models.

Multimodal Model Diffing for Feature Discovery and Control A “diff” tool for AI: Finding behavioral differences in new models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.894475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.894475Z digest=sha256:a00919b5a54b7a0a9f6a8d8f96aa65f4dfded31e405fe99b32dd2fd356df324f

Observation 7bdee4c9-2165-49f5-a2c8-8a114e2a8644 · outbound

This paper cites Bridging the VLM and mech interp communities for multimodal interpretability.

Multimodal Model Diffing for Feature Discovery and Control Bridging the VLM and mech interp communities for multimodal interpretability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.900408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.900408Z digest=sha256:269958385c3209b0564d6b178bd1e21b35d74074f57f8a1c9bb188ffd3d57e0a

Observation 0a1b3935-db61-4b7d-b663-d9f6cdf18b9d · outbound

This paper cites Steering CLIP's vision transformer with sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Steering CLIP's vision transformer with sparse autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.905911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.905911Z digest=sha256:64e01f9e38a23bd83b8e206077f67263c3f8d63456589a8ea8d5869e9045f373

Observation 02d7cbe7-c662-4f20-a2f1-8f252dca6f28 · outbound

This paper cites Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video.

Multimodal Model Diffing for Feature Discovery and Control Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.911157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.911157Z digest=sha256:875aa8f09ac31ad1488514ae1ab10d664fee11a7533d4ac19d59c7f73abe3bb3

Observation 4e56b9bf-2caf-4a64-8b63-4fffadd4fa48 · outbound

This paper cites Analyzing Finetuning Representation Shift for Multimodal LLMs Steering.

Multimodal Model Diffing for Feature Discovery and Control Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.916191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.916191Z digest=sha256:7bc27689412efc76ed33b8d46b97811c80642d545d457d78dba9dcb5bed53f7e

Observation bb729be5-e31f-4c6f-95d7-d9726ae0138e · outbound

This paper cites Saes (usually) transfer between base and chat models.

Multimodal Model Diffing for Feature Discovery and Control Saes (usually) transfer between base and chat models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.921070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.921070Z digest=sha256:e08c94166947fca6b784576927bb4cee2a470e8f8b622e77ea81c1a5f4843c79

Observation 16f919f1-5867-4ed1-b1a8-49dea9800dbf · outbound

This paper cites Similarity of neural network representations revisited.

Multimodal Model Diffing for Feature Discovery and Control Similarity of neural network representations revisited

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.926358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.926358Z digest=sha256:97bcb54347dbfbc5b0abb5d80493aec250f58f78401b1987e83e2f0c342dae98

Observation 07a5bec2-b757-4592-8d0b-4eccf85b6d89 · outbound

This paper cites Sakla, and Kowshik Thopalli.

Multimodal Model Diffing for Feature Discovery and Control Sakla, and Kowshik Thopalli

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.931559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.931559Z digest=sha256:262f0765366954413ca5060c3eb5a54bbede3e8663ae2357bf910ca74e815dcf

Observation 62157067-b288-4ad7-97cb-c5b4e1633c98 · outbound

This paper cites Understanding image representations by measuring their equivariance and equivalence.

Multimodal Model Diffing for Feature Discovery and Control Understanding image representations by measuring their equivariance and equivalence

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.936739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.936739Z digest=sha256:67684cedde1e4b2fdfbe25fac35006c2ec768ea0315079eb5b2c46c71d2e1d6c

Observation ef4f6fd3-9c05-4ffc-8be7-8a071882afe6 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.941614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.941614Z digest=sha256:1c478c35af7c65e6efc82b5f6a1a5b8051b4f74a8a31ddd9772ad60f8594b6f1

Observation ae34dab8-af93-4449-8e3a-061503ffd727 · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

Multimodal Model Diffing for Feature Discovery and Control Inference- time intervention: Eliciting truthful answers from a language model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.946783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.946783Z digest=sha256:be40034b00888086da0bece714934bd3349b14e02cabd592f8044150dfc9cbad

Observation f66078d7-2322-48be-8ac4-5c1c808c4191 · outbound

This paper cites Images are Achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control Images are Achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.969271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:55.951510Z digest=sha256:43daf86da926ca1965c408b1262d819618d260938762b88a105e44493df13fea

Observation 7c5cf49e-69d0-46fc-b00c-d062e287ae2c · outbound

This paper cites Convergent Learning: Do different neural networks learn the same representations?.

Multimodal Model Diffing for Feature Discovery and Control Convergent Learning: Do different neural networks learn the same representations?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.956223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.956223Z digest=sha256:a8637765d424ccdbf42e3f92c8cd548685a6076a77acb576c00307a871176dd6

Observation b362ad44-4a2b-4a9d-9394-b8618f4cc42d · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Multimodal Model Diffing for Feature Discovery and Control Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.961255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.961255Z digest=sha256:cb14683331001aeda86dae5be6d3110b386ccc7f116a5b7714451bbd6ee534f9

Observation f5eddaa7-8d8b-40f1-9231-425afcca5f9e · outbound

This paper cites Sparse autoencoders reveal selective remapping of visual concepts during adaptation.

Multimodal Model Diffing for Feature Discovery and Control Sparse autoencoders reveal selective remapping of visual concepts during adaptation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.967019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.967019Z digest=sha256:178ac38f966c6254451b66f20359927a6a55f29c23f5fcdfc50d6f37d7a2d7c6

Observation 74cb08c5-3e38-4f66-b5a3-cf3e6d9061f1 · outbound

This paper cites A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models.

Multimodal Model Diffing for Feature Discovery and Control A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.972218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.972218Z digest=sha256:80564e20c4ed23ee7f391dc4a7486036ecdd1f69e8690066f49429f67f746da9

Observation 50c39459-8359-4d8a-8f3b-06a5764fccdb · outbound

This paper cites Sparse crosscoders for cross-layer features and model diffing, October 25 2024.

Multimodal Model Diffing for Feature Discovery and Control Sparse crosscoders for cross-layer features and model diffing, October 25 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.952495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:55.977277Z digest=sha256:85809ad10b986b2a8c814a0973608b16323bdfb6647095e72726258d1674ec4d

Observation e242e417-d6d1-466c-8c67-75d87eb17f48 · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 2023.

Multimodal Model Diffing for Feature Discovery and Control Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.934987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:55.981926Z digest=sha256:d20acda0bf4dec16c5e589266dd3cc0eb360a7936fe76d8a33d2683eacb8f330

Observation 428d81df-84f7-4297-95c6-84ba78dfca31 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Multimodal Model Diffing for Feature Discovery and Control Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.986160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.986160Z digest=sha256:b646d4941b382f8d22aeedaaf869169ded9abb932093201532c4584e82df0dad

Observation 159ce224-eb30-4005-bd03-3e7034b33e31 · outbound

This paper cites Improved baselines with visual instruction tuning.

Multimodal Model Diffing for Feature Discovery and Control Improved baselines with visual instruction tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.990284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.990284Z digest=sha256:61db913cbbd93d5bc511b58639b1cd30b16c6f09156bb53f0e9d2e5d67d171bc

Observation a8f7a635-04af-41c1-9058-446d263e58fb · outbound

This paper cites MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.897139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:55.994689Z digest=sha256:cfb446217a379ad07be988a5acc56e7c38469f79cfb7071bcd942ee2d36e8779

Observation 4cc32df0-f6f3-4a06-9080-8ad98d0fa6c4 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Multimodal Model Diffing for Feature Discovery and Control OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.999133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.999133Z digest=sha256:08066b8379993f1f92533497f84f89739722a9738aefb932098953308f303445

Observation 2db28554-8e14-41be-8bc8-fc9bd7333573 · outbound

This paper cites Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller.

Multimodal Model Diffing for Feature Discovery and Control Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.880397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.003768Z digest=sha256:e0281dc7b1fcb018d6f44fe5c28aaf9356e151a08420394c8a8b9e2aa461fc18

Observation 1012ade3-0d1d-4240-a17c-094795998f39 · outbound

This paper cites Locating and editing factual associations in GPT.

Multimodal Model Diffing for Feature Discovery and Control Locating and editing factual associations in GPT

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.008128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.008128Z digest=sha256:567bce0e7f5b05f214f9d4e8e4a141f09389026086f432cc1fef96509e96f108

Observation 464d7b55-d456-457f-86c9-7b103017259e · outbound

This paper cites Robustly identifying concepts introduced during chat fine-tuning using crosscoders.arXiv preprint arXiv:2504.02922, 2025.

Multimodal Model Diffing for Feature Discovery and Control Robustly identifying concepts introduced during chat fine-tuning using crosscoders.arXiv preprint arXiv:2504.02922, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.012960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.012960Z digest=sha256:f9c634beddb52eed7754348d95941d3c047422415f4c2709abfc1d871e68586c

Observation cc5bc455-3eba-42e8-901f-b572d332ca4a · outbound

This paper cites What we learned trying to diff base and chat models (and why it matters).LessWrong, 2025.

Multimodal Model Diffing for Feature Discovery and Control What we learned trying to diff base and chat models (and why it matters).LessWrong, 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.853986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.017719Z digest=sha256:671b9d57f069ef689d15cef5e5dc5af9f4a2f2a58475d25494e22849357e92ee

Observation 8451658b-066d-4558-848c-483cfff8b398 · outbound

This paper cites Insights on crosscoder model diffing.

Multimodal Model Diffing for Feature Discovery and Control Insights on crosscoder model diffing

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.837250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.022501Z digest=sha256:b1413b81b6c9ac8c9717bd8e26490983f75a498ad25df7d47248995d9b178737

Observation 5db2cfa0-3178-48d0-9c2f-245259e280e8 · outbound

This paper cites Attribution patching: Activation patching at industrial scale.

Multimodal Model Diffing for Feature Discovery and Control Attribution patching: Activation patching at industrial scale

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.818307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.027425Z digest=sha256:7f898fa70397dfa9175404eeb92ba0dbf970679da2c29d2de2f104967fefa98c

Observation cdd46c79-02f1-4c4e-86bb-b856ba020424 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Multimodal Model Diffing for Feature Discovery and Control Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.032176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.032176Z digest=sha256:544ad2f6787879fb0cdc5ab1a16115455bfbc84d1cd6ee2a72f3b6bc85d8b4f5

Observation 623a2370-84b8-4f4a-bf0b-0b8c4a7bd95c · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Steering Language Model Refusal with Sparse Autoencoders

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.038373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.038373Z digest=sha256:9763f9c143020d1f7841ce11571e5b5a155b09d0b2b412c5400fb60a4f20d1c6

Observation bbc722ed-96ce-4a69-89d5-624cf509d697 · outbound

This paper cites Zoom in: An introduction to circuits.Distill, 2020.

Multimodal Model Diffing for Feature Discovery and Control Zoom in: An introduction to circuits.Distill, 2020

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.044310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.044310Z digest=sha256:643f47ed3577879f98b36f3b1c4f17aea62d33ced5ed3744f7b70faf3ded5971

Observation 058d2038-410f-42cd-9c2b-798bf255d421 · outbound

This paper cites Visualizing representations: Deep learning and human beings.

Multimodal Model Diffing for Feature Discovery and Control Visualizing representations: Deep learning and human beings

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.798542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.049202Z digest=sha256:b3ababbb9bdcc836fb7ca8f9e86ee2f62f747d26962ddc50d8912ac8faa61cf2

Observation 60c5603e-ea1d-40e2-b130-362652456e2b · outbound

This paper cites Probing the representational power of sparse autoencoders in vision models.

Multimodal Model Diffing for Feature Discovery and Control Probing the representational power of sparse autoencoders in vision models

Reference 65

Resolution
verified exact
raw_fallback, observed 2026-08-11T04:17:56.747243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.054083Z digest=sha256:8af565e2a5d0b9c0485035b26b2edad1420170a83f4591e2999dfae25eaf0b3f

Observation 5b9793e9-4c4a-4e7c-862e-920e5d638bbe · outbound

This paper cites Gpt-4o-mini: Advancing cost-efficient intelligence.

Multimodal Model Diffing for Feature Discovery and Control Gpt-4o-mini: Advancing cost-efficient intelligence

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.782692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.058945Z digest=sha256:94479e5cdfc6121c56f8f7222f82ffc4a5e70c05188e4a8d4981ca1fcf6e36eb

Observation 377960dc-cb41-4e38-a7f4-d2932749b680 · outbound

This paper cites Sparse autoencoders learn monosemantic features in vision-language models.

Multimodal Model Diffing for Feature Discovery and Control Sparse autoencoders learn monosemantic features in vision-language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.063731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.063731Z digest=sha256:e80b66cfdc4743cb506f118fd84e651eeddb5f53dc97faa590e83867090dafd7

Observation 1c619101-97ef-42f1-aae8-e70c8d36057a · outbound

This paper cites Towards vision-language mechanistic interpretability: A causal tracing tool for blip.

Multimodal Model Diffing for Feature Discovery and Control Towards vision-language mechanistic interpretability: A causal tracing tool for blip

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.766065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.068462Z digest=sha256:f0a6f8fbb7245ac75f744adb851bc3d9ad03491e93b497a40201653f981d04ef

Observation 94b2dead-18bc-4a07-af33-a10015ae3873 · outbound

This paper cites Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal.

Multimodal Model Diffing for Feature Discovery and Control Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.073278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.073278Z digest=sha256:e16fb650fec91289c89cd5bd258ea27e89951e4a3d279b0604e38127fd0110d7

Observation eacb6670-9287-4508-b7cc-10e65ab0b142 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

Multimodal Model Diffing for Feature Discovery and Control Visual adversarial examples jailbreak aligned large language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.750674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.078497Z digest=sha256:28ce310c479307a52ecd29dd6521d89f3d356e15c474ee713c8099eae7f4584a

Observation d5566a4c-bbc4-4029-82c9-0268f9e08351 · outbound

This paper cites Qwen-Scope: An open sparse autoencoder suite for the Qwen model family.

Multimodal Model Diffing for Feature Discovery and Control Qwen-Scope: An open sparse autoencoder suite for the Qwen model family

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.733675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.083263Z digest=sha256:af2a0c8001082be90489a8975bd73da6ae98975906fcb7ac1342b1639c902a3d

Observation fe447d69-874d-4cbd-b5a0-ea082f32f82e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Multimodal Model Diffing for Feature Discovery and Control Learning transferable visual models from natural language supervision

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.088107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.088107Z digest=sha256:0c29501a44815ea0856e7e9cce3a6640e74923260779137850f9b8748dee834e

Observation eadbd835-4e5f-45e3-8fea-de3c5e559122 · outbound

This paper cites Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.092721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.092721Z digest=sha256:1d1197821c0c6840fe6d6318ba319b0b84f35e7634e6e20a19ac5f521f049d42

Observation 4641de98-7e7b-4fc5-a4ab-6302312a8412 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Multimodal Model Diffing for Feature Discovery and Control Steering Llama 2 via Contrastive Activation Addition

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.097463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.097463Z digest=sha256:a8fe4931b65b76825371b3ca54e56ec252b393ec4784f1a92a99d6be55447fc3

Observation 8c892395-c64d-4597-88bb-e01c606b2917 · outbound

This paper cites Multi- modal neurons in pretrained text-only transformers.

Multimodal Model Diffing for Feature Discovery and Control Multi- modal neurons in pretrained text-only transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.705410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.102224Z digest=sha256:90ac1cb5ce8f2546e024656e7219ae503bba5ebb62b7fdd23c0196c6ebc93a78

Observation f298f62e-2df4-4676-8705-f70f9d84fe17 · outbound

This paper cites SteerVLM: Robust model control through lightweight activation steering for vision language models.

Multimodal Model Diffing for Feature Discovery and Control SteerVLM: Robust model control through lightweight activation steering for vision language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.688536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.106645Z digest=sha256:c296f3ef22847fb99d428d33d5556e73cd6dc4d558d37d98c79dfdd51d1eba94

Observation 830f7044-5311-48c7-b4a4-e492efe747ce · outbound

This paper cites LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models.

Multimodal Model Diffing for Feature Discovery and Control LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.111721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.111721Z digest=sha256:edde199d2c4518d322daff10fb9d462c25b0cc31112d26b67133a59b7f2d2e48

Observation ea2aabab-eee9-460d-ba2a-2b5538ff20e5 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Multimodal Model Diffing for Feature Discovery and Control PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.116317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.116317Z digest=sha256:237dac6e95fed2a4107d27b9b214172c22ba4861ee6153ded6b01540a8bbae8a

Observation 652c6242-0c83-45b3-84c8-a17a7d903b88 · outbound

This paper cites Daniel Freeman, Theodore R.

Multimodal Model Diffing for Feature Discovery and Control Daniel Freeman, Theodore R

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.122564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.122564Z digest=sha256:3fe99cef14c57ea6469efd28a8d63631eb36cf6ec74695c5cb0515b09e2f665f

Observation ad666547-3921-415d-9f0d-0f41acd4ac5e · outbound

This paper cites Li, Arnab Sen Sharma, Aaron Mueller, Byron C.

Multimodal Model Diffing for Feature Discovery and Control Li, Arnab Sen Sharma, Aaron Mueller, Byron C

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.659875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.127927Z digest=sha256:6f36854723062d9c4df17e3d626acf060805ac6836e34f1a5b34f21724f48c72

Observation d37558b3-dc39-493d-bf70-b84cb8eedef5 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Multimodal Model Diffing for Feature Discovery and Control Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.132540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.132540Z digest=sha256:c6fe2ce719e7a6d36f0828cebbaa8de12a913d40b2b571ee7ba9efb8dddf35eb

Observation 4f30b446-4ab4-4a1a-a778-10e1ba21a3ed · outbound

This paper cites Steering Language Models With Activation Engineering.

Multimodal Model Diffing for Feature Discovery and Control Steering Language Models With Activation Engineering

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.137372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.137372Z digest=sha256:140df2f24c83da88fdd19e330f7a724927c39779911dbc11a093f4ac20a2ae9b

Observation a5e3ee40-f7fd-46f4-9da1-684f3115af2c · outbound

This paper cites Too late to recall: The two-hop problem in multimodal knowledge retrieval.

Multimodal Model Diffing for Feature Discovery and Control Too late to recall: The two-hop problem in multimodal knowledge retrieval

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.628662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.142365Z digest=sha256:1d10b049a0c94cf00d1032251cdfb8aa2695c91747be5df1ccd9ba43bb591519

Observation e97f989c-c33a-4426-8979-5210adbd71a5 · outbound

This paper cites How Visual Representations Map to Language Feature Space in Multimodal LLMs.

Multimodal Model Diffing for Feature Discovery and Control How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.147155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.147155Z digest=sha256:f6efa46188e0710d8e5d6eec706503efe60c3564f3e03018044f8b5a662fe1a7

Observation 9ddfa148-7c15-4249-9c65-72225f1a76bd · outbound

This paper cites Steering away from harm: An adaptive approach to defending vision language model against jailbreaks.

Multimodal Model Diffing for Feature Discovery and Control Steering away from harm: An adaptive approach to defending vision language model against jailbreaks

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.611779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.152172Z digest=sha256:cebd4cb7b003bd17e519f4ed6bb7a130ad557be7d0515411a90d3418105e163b

Observation 3ad407f8-f160-46c8-9936-e40a5904103c · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Multimodal Model Diffing for Feature Discovery and Control InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.156829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.156829Z digest=sha256:e875e2436cc600d7c22444e9ea6fac2d48c169295bdde9e065596fe84b17759b

Observation 3b14a2e9-6053-4aa2-9b6a-279769fe4d46 · outbound

This paper cites AdaShield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting.

Multimodal Model Diffing for Feature Discovery and Control AdaShield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.595098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.162853Z digest=sha256:4cd8f192a598147daed4dc9b25e5279092f20734baaf222778d2066c80f80bbb

Observation 47a8435a-2c41-4f4a-81e8-3c19919d4f69 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.167527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.167527Z digest=sha256:25e1a36107eff7039f8caa5366892721fa9ecc2cd0996f53c88bdcd0eaa97d44

Observation 95f2b71b-43ef-4498-9931-0584ebf2d855 · outbound

This paper cites Qwen3 Technical Report.

Multimodal Model Diffing for Feature Discovery and Control Qwen3 Technical Report

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.172320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.172320Z digest=sha256:255669816dbdde6e7bd7c87e0ad28b95e86baeebea3353b8e3fa189fa1b1ed25

Observation 581e9dcf-2c1f-4833-b068-0610fedbe8bf · outbound

This paper cites SafeSteer: Adaptive subspace steering for efficient jailbreak defense in vision-language models.arXiv preprint arXiv:2509.21400, 2025.

Multimodal Model Diffing for Feature Discovery and Control SafeSteer: Adaptive subspace steering for efficient jailbreak defense in vision-language models.arXiv preprint arXiv:2509.21400, 2025

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.177100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.177100Z digest=sha256:96fca674dd689fb6da75823f9f41f36653b7ea68739831bd6c609e8f093f2da1

Observation c093c9d3-d5cc-4865-8131-636221f4691f · outbound

This paper cites Sigmoid loss for language image pre-training.

Multimodal Model Diffing for Feature Discovery and Control Sigmoid loss for language image pre-training

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.181619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.181619Z digest=sha256:438a0d6d37539763069fa87adb467d38eb50c047ed128883ba72e2057b686da7

Observation 112922d6-827d-42aa-a3d2-3aea70380e8f · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Multimodal Model Diffing for Feature Discovery and Control Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.186065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.186065Z digest=sha256:4ed78edb4d3872984c0513463a9d48a0d79f9e30a5fc40f2d876d93f6d260017

Observation f2b68228-1a9b-4451-b843-561f9a915118 · outbound

This paper cites Cross-modal information flow in multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control Cross-modal information flow in multimodal large language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.566959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.190937Z digest=sha256:b8a512a918d38a4534b3600a619a0ad2063b4bb7bf93fe791ffa07545bf9c45a

Observation 2619e432-0229-4ce4-9384-2c2445bcae20 · outbound

This paper cites Multimodal situational safety.

Multimodal Model Diffing for Feature Discovery and Control Multimodal situational safety

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.549778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.195809Z digest=sha256:c9dc85d307572e9bf4e93f22c9fef25b9d836c7fe60edf3b054d6719c8014ed6

Observation e7d95dd4-96b3-42a8-bfa2-b16e8a5e4b3e · outbound

This paper cites Relocated.

Multimodal Model Diffing for Feature Discovery and Control Relocated

Reference 95

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:17:57.530047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.201760Z digest=sha256:4965635dc949293a8058ed40174b958e4df1d377869c61e1842679e04f759217

Observation dd40ce95-4dfb-4b0e-bb9f-d4bfd52a465d · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.512239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.208372Z digest=sha256:61448ef5b313ee517fe1ca14725d77ee2675e77abbb1634d2085ca12619abbbb

Observation 7bbbfb87-1a52-4a94-8530-dc262d50d734 · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.494613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.214057Z digest=sha256:b11698ecef75693c7b230ba7ad773e9eaccda0621348be33dab04a61ae884e86

Observation 418b9983-293a-4b05-83aa-d201d05c83b7 · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.478120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.219618Z digest=sha256:7e747a2e36aeef173864fa8573be48962ab81b8cd0f6b8415ced289e5bb616b5

Observation 4447cbf9-6ef0-4020-9860-8acf0b877876 · outbound

This paper cites this neuron activates for.

Multimodal Model Diffing for Feature Discovery and Control this neuron activates for

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.462367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:17:56.224644Z digest=sha256:e071f2468a6a5c5ba924069f6b5597ff90e3785722463c5054069cccf54487e6

Pith citing papers

No inbound Pith citation observations are available.